Search++ (\W)'(\w) regex replace failure
-
-
I didn’t experiment with adding or subtracting any plugins (other than having already added Search++) in the portable version yet. but I did just open the portable again, this time via command line and
-multiInst, per @Alan-Kilborn’s suggestion (thanks, Alan), while my installed version was already open. I then retried my regex, but this time got anomalous results, such as:this u’dos I am that this ’Round and ’round thatI then undid all replacements, and redid them, with all results correct the second time around.
Quit the portable copy, then reopened it (as
-multiInst) again, then redid the regex, with anomalies:this ’Cos I am that this ’Round and-’-ound thatQuit the portable copy again, then reopened it (as
-multiInst) again, then redid the regex again, with slightly different anomalies:this ’Cos I am that this -’-ound and ’round thatTalk about bizarre. But it seems we can rule out conflicts from the other plugins I use in my installed version, unless there’s some way for them (or something else) to cross-contaminate the process in active memory due to the installed version still being open.
[…shuts down both portable and installed versions; reopens portable version; retries regex…]
Nope. Still had anomalous results:
this ’Cos I am that this -’-ound and ’round thatI’m out of guesses. It’s hit or miss whether or where or how it happens.
-
I’m out of guesses. It’s hit or miss whether or where or how it happens.
Thank you for the reports. I don’t know the cause yet either, but I now think it is very likely that it is a fault in my code which I just haven’t encountered yet.
I notice that the literal
’is always correct; when something is wrong, it’s the\1and\2substitutions. I have a hunch that if you were to show line endings, in the ones where a spurious character appears before a’at the beginning of a line the preceding line ending would have lost its LF and become just a lone CR.I will work on it. Thank you again for your patience and this information. The reports that it occurred in a clean, portable version are very helpful in that they tell me what it isn’t; with something I can’t readily reproduce, that’s valuable knowledge.
-
I notice that the literal
’is always correctAh, but don’t forget in my initial post, I was sometimes getting question marks (
?) instead of’. When I first saw that, I suspected that maybe Search++ was failing to make the conversion back from UTF to ANSI, since I believe it was previously stated that text is treated internally as UTF, but I was actually working on ANSI text at the time, and a question mark is what I usually get when accidentally trying to paste a UTF character in an ANSI file. But, I guess that’s not the case. -
I have a hunch that if you were to show line endings, in the ones where a spurious character appears before a
’at the beginning of a line the preceding line ending would have lost its LF and become just a lone CR.Here ya go… This is what happened when I just now tried it:

Note that only one CRLF became just LF (line 8), but also that in the case of what you see here as line 5, it’s actually a combination of what were lines 5 and 6 from the pre-replacement text, so an entire CRLF is now missing. I don’t particularly think that’s related to the fact that I had View > Show Symbol > Show End of Line enabled, but what do I know?
-
M-Andre-Z-Eckenrode said:
Note that only one CRLF became just LF (line 8)
Correction: Only one CRLF became just CR.
-
Hello, @m-andre-z-eckenrode, @alan-kilborn, @coises, @mpheath, @thomas-knoefel and All,
Just an observation : if you modify the proposed replacement :
FIND
(\W)'(\w)REPLACE
\1’\2by this one :
FIND
(?<=\W)'(?=\w)REPLACE
’No problem occurs for a step by step replacement sequence or a
Replace Alloperation for, either, anANSIorUTF-8encoded file. I verified this assumption with four regex engines :-
The native
Search > Replacedialog of N++ -
The
Plugins > Column++ > Search...dialog -
The
Plugins > Search++ > Search...dialog -
The
Plugins > MultiReplace > MultiReplace...dialog
May be, this result will help to resolve possible bugs, with the other regex syntaxes
Best Regards,
guy038
-
-
I’ve been looking, but so far I have not found a potential cause for the results you obtained.
@guy038’s post just led me to think of something.
The debug information in the post where you asked about the failure to replace in Notepad++ showed that you have Columns++ installed.
The search and replace code Search++ uses is derived from what I used in Columns++, but I have made some changes. Would you try and see if you can get the search in Columns++ to misbehave under the same circumstances as the one in Search++? (It’s fine to use the installed version of Notepad++ for this test.)
Knowing whether or not the fault also appears in Columns++ could help me narrow the scope of what I could be missing.
As always, I understand if you don’t have the time for this sort of thing. Thank you for your observations so far, and many thanks in advance if you are willing to do this additional test.
-
Hi, @m-andre-z-eckenrode, @alan-kilborn, @coises, @mpheath, @thomas-knoefel and All,
@coises, personally, I confirm that the step by step replacement with the regex pattern :
-
FIND
(\W)'(\w) -
REPLACE
\1’\2
Works correctly with both your two plugins
Columns++andSearch++and with theMultiReplaceplugin as well, when using the N++v8.9release !The bug seems to occur only with native N++ regex engine !
BR
guy038
-
-
Would you try and see if you can get the search in Columns++ to misbehave under the same circumstances as the one in Search++?
In installed NPP, using Columns++ search in regex mode with a rectangular selection of multiple copies of my previously shown example text, I find that only the instances of “
'r” (preceded by a regular space, NOT by any EOL) are matched, and are all replaced correctly with no anomalous characters. To be clear, this, again, is the starting text:this 'Cos I am that this 'Round and 'round that this 'Cos I am that this 'Round and 'round thatAnd resulting text:
this 'Cos I am that this 'Round and ’round that this 'Cos I am that this 'Round and ’round that -
-
Hello, @m-andre-z-eckenrode, @coises and All,
@m-andre-z-eckenrode, you said :
I’m confused. You’re saying that a native N++ regex replacement, using my example text and regex strings, results in anomalous characters for you? I haven’t seen that at all.
I didn’t say that anormalous characters appear. I just said that no replacement occurs at all !
Indeed, with the native regex N++ replacement, and the text below :
this 'Cos I am that this 'Round and 'round thatand with the following REPLACEMENT :
-
FIND
(\W)'(\w) -
REPLACE
\1’\2
The first two replacements do not occur ! Only the last one will change the normal quote
'into the’character (\x{2019}RIGHT SINGLE QUOTATION MARK )IMPORTANT : This test was done with N++
v8.9. May be, the lastv8.9.7release would produce a different result !Best Regards,
guy038
-
-
using Columns++ search in regex mode with a rectangular selection of multiple copies of my previously shown example text, I find that only the instances of “ 'r” (preceded by a regular space, NOT by any EOL) are matched
I’m sorry… I didn’t think to mention that quirk. If you start with a rectangular selection, Columns++ searches each row independently, so the ones with
'at the beginning of a line would fail to match.If Auto set is checked (the default) in the Search in indicated region dialog and no region is already indicated, when you start a search with nothing selected the indicated region is set to include the whole document. (You can also select the entire document, or make any multiple line selection or multiple selection, before you search.¹) This is different from a rectangular selection converted to an indicated region: rectangular selections are multiple selections which never include line endings.
(Yet another complication is that Scintilla does not reliably restore indicators (marked text) on undo. I think this only affects cases where a replacement begins at the first character of a contiguous span of marked text — Scintilla restores the text, but not the marker, on undo.)
So it is expected that only occurrences within a single line would match if you start with a rectangular selection. If you start with the entire document selected, or with no selection at all, it should match all the same occurrences as Search++ (with no selection or marked region).
¹ The rules are complicated. The confusing nature of the “indicated region” was one of the reasons I wanted to split Search++ entirely instead of doing more work on the search in Columns++. I didn’t want to change the concept for people who are already used to Columns++ search, but outside of the specific case of searching in column selections, I think it is more confusing than it needs to be.
-
I didn’t say that anormalous characters appear. I just said that no replacement occurs at all !
You might be confusing @m-andre-z-eckenrode’s topic (\W)'(\w) Regex replace failure, about Notepad++, and this topic, about Search++, with the same search.
He was surprised that Notepad++ found the text but would not replace it. The cause is how Notepad++ treats CRLF pairs combined with how it decides whether Replace should replace or find next — as I wrote there, whether it is a bug is, I suppose, a matter of opinion.
Search++ has a different potentially user-unfriendly behavior: with different find or replace expressions (any regular expression that can match starting with an LF, combined with a replacement string that doesn’t copy the first matched character as the first character of the replacement) a user could invisibly change line endings from CRLF to just CR (though if the user understands the regular expression and replacement entered, it would be expected).
Notepad++ has the same behavior as Search++ when using Replace All (including that it can replace the LF in a CRLF pair with something else); it’s just step-by-step replace that can find but fail to replace (for the same reason that
\Kdoesn’t work in step-by-step native searches).
In this topic, the problem is that Search++ is, apparently randomly, replacing with garbage instead of the correct replacement string.
I have not yet found a possible cause. It does not happen on my machine.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login