Community
    • Login

    Search++: A work in progress

    Scheduled Pinned Locked Moved Notepad++ & Plugin Development
    208 Posts 12 Posters 62.0k Views 3 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • guy038G Online
      guy038
      last edited by guy038

      Hi, @coises,

      I’ve just tested your new v.0.7.1 release !


      Let’s suppose that we marked the occurrences of the word test in the text below, that you’ll paste in a new tab :

      
      bla
      blah
      This is the test to do
      bla
      blah
      test
      bla
      blah
      

      And that we checked the Bookmark lines when marking text (Shift: configure option, in the Tools menu

      => The line 4 and 7 are bookmarked and the two occurrences of the word test are marked

      In an old post, I said that, in order to get the complementary regex, we had to use the ^((?!test).)*$ regex as below :

      47217b17-e9ac-499b-9779-b14417deb58c-image.jpeg

      Luckily, there’s no need to find out this complementary regex, anymore :

      • First, mark all the occurrences of the word test.

      d77b5da7-0dfe-4d8f-a5a4-d8650672eb92-image.jpeg

      • Right-click within the bookmark margin of N++ and choose the Inverse bookmarks option.

      022b778b-1455-4dd8-85be-3938d80c0dcf-image.jpeg

      • Then, go to the Synchronize marks and bookmarks option of the Tools menu and choose the Mark all text in bookmarked lines and clear other marks option

      => You should get this snapshot :

      b7c6ec3a-3e72-4e1a-b229-52ee2bc7a638-image.jpeg

      Which correctly bookmarked all the lines not containing the word test and then marked all contents of the bookmarked lines, including the EOLs, as well !


      I tested, both :

      • The Bookmark lines with marked text, and clear other bookmarks option VS the Add bookmarks to lines with marked text option

      • The Mark all text in bookmarked lines, and clear other marks option VS the Add marks to all text in bookmarked lines option

      And, if I understand you correctly, you mean that the Synchronize marks and bookmarks sub-menu, in your future release, will look as below :

      Bookmark lines with marked text, and clear other bookmarks
      Bookmark visible lines, and clear other bookmarks
      Mark all text in bookmarked lines, and clear other marks
      
      Add bookmarks to lines with marked text
      Add bookmarks to visible lines
      Add marks to all text in bookmarked lines
      

      Yes, that’s totally coherent : The first three items do a clear operation whereas the last three items simply add bookmarks / marks to the existing layout !

      Best Regards

      guy038

      P.S. :

      Just a random question about regex. In the sample of text above :

      • If I try to mark the \z regex with your plugin, I do get the bookmarkred line 10, which is the final empty line of the text

      • If I mark all the occurrences of the (?-s)^((?!test).)*\R regex => All lines, but lines 4, 7, 10, are bookmarked and text with EOLs, in lines 1, 2, 3, 5, 6, 8, 9 is marked

      However, marking all occurrences of the ^(?-s)((?!test).)*\R|\z regex do not bookmark the final line 10, nor the (?-s)^((?!test).)*$\R? regex and nor the (?-s)^((?!test).)*(\R|\z) regex !

      Not important, though. Do you have an explanation ?

      CoisesC 1 Reply Last reply Reply Quote 1
      • CoisesC Online
        Coises @guy038
        last edited by Coises

        @guy038 said:

        Do you have an explanation ?

        This is exceedingly odd. Thank you for identifying and describing it clearly. I do not have an explanation, but I think it has to be a bug (and not just a quirk of regex). Here’s why.

        We can’t compare the same test in native Notepad++ search, because Mark doesn’t count or bookmark null matches. (The same with Count and Find All: in Notepad++ native search, they ignore null matches.)

        However, single-step Find does identify null matches in both engines. In Notepad++ native search using Find Next, your expression (?-s)^((?!test).)*\R|\z identifies eight matches. The last is at the end of the file and shows the ^ zero length match message.

        In Search++ Regex with Find, it only finds seven.

        What is even more peculiar is if you switch to the ICU engine and Mark the same expressions. In ICU as in Regex, \z alone does as you would expect, bookmarking the last line and showing the message “Found 1 match (0 marked, 1 null).” ICU also does the same as Regex with (?-s)^((?!test).)*\R, marking and bookmarking seven lines, exactly as you would expect.

        Adding the |\z to the end changes nothing in Regex, though you would expect to see “Found 8 matches (7 marked, 1 null)” and a bookmark on the last line. In ICU it not only doesn’t bookmark the last line, it finds and bookmarks nine matches, marking the LF, but not the CR or the rest of the line, on the lines containing test!

        As far as I can tell (I might not have tested every possible combination yet), both Regex and ICU are self-consistent in that they Find and Count and Find All the same things they mark/bookmark.

        Why adding |\z causes such strange behavior — in ICU even leading to new matches in the middle of the file — as yet I have no clue.

        CoisesC 1 Reply Last reply Reply Quote 1
        • CoisesC Online
          Coises @Coises
          last edited by

          @guy038 said:

          Do you have an explanation ?

          and I wrote:

          I think it has to be a bug

          It is, and it’s a fairly simple error in my logic. If a match ends at the end of the file (or the end of a selection or a span of marked text, if you’re searching within one of those), the search ends (or moves on to the next selection or span of marked text) without checking whether the expression can match null there. It will be fixed in the next release.

          Why adding |\z causes such strange behavior — in ICU even leading to new matches in the middle of the file — as yet I have no clue.

          The code supporting ICU search has the same logic error as the Regex search does, but there’s an additional oddity which seems to be built into ICU’s regular expression search. As best I can tell, it treats the circumflex (^) differently when it is the first element of the only alternative and in all other contexts. As an example, see the difference between marking ^\n (which matches nothing in Guy’s test text) and ^\n|Q (which matches the line feed at the end of each of the first nine lines).

          Note: I suggested using Mark to show this because Scintilla will not allow a selection to encompass only part of a line ending, so it’s less clear what is happening with Select or stepwise Find commands. The Search++ results list doesn’t have the option to show line endings — perhaps it should! — so Find All is also unclear.

          I’m not sure if this is expected behavior. ICU describes the circumflex as “Match at the beginning of a line” but, as far as I can find, says nothing about it behaving differently in different positions or expressions.

          1 Reply Last reply Reply Quote 1
          • guy038G Online
            guy038
            last edited by guy038

            Hi, @coises, @thomas-knoefel and All,

            FIRST test :

            The only way to bookmark the last empty line 10, that I found, is to use the regex ^((?!test).)*$ again my example text, below :

            
            bla
            blah
            This is the test to do
            bla
            blah
            test
            bla
            blah
            

            Some observations :

            • Using Search++, in mode Regex, this regex counts 8 matches, whose line 1 with ONLY a small blue triangle at bottom of line 1 and the last empty line 10 with the calltip ^ zero length match

            • Using Columns++, this regex counts also 8 matches, but, this time, the first line returns, instead, the calltip ^ zero length match

            • Using Search++, in mode ICU this regex counts 7 matches, whose line 1 with ONLY a small blue triangle at bottom of line 1. So the line 10 is not a match


            SECOND test :

            Paste the text below in a new tab and search, in regex mode, for the regex ^abc

            abc
            
            abcabc    0011
            abc       0012
            
            abc…abc   0085
            
abc       2028
            
abc       2029
            abc         After the above \r\n
            ====================================
            	abc     0009
             abc        0020
             abc     00A0
            ­abc      00AD
             abc     2000
             abc      200A
             abc    202F
             abc     205F
             abc     3000
            ​abc     200B
            ‌abc     200C
            ‍abc      200D
            ⁠abc       2060
            abc   FEFF
            abc      FFFA
            abc      FFFB
            abc      FFFC
            

            May be, after pasting, you’ll have to change the line 2 which needs to be an LF char ONLY and line 5 which needs to be an CR char ONLY !

            • With N++ and MultiReplace the ^abc regex counts 8 matches. So a beginning of line is seen :

              • At the very beginning of the file

              • After a LF character

              • After a FF character

              • After a CR character

              • After a NEL character

              • After a LS character

              • After a PS character

              • After a \r\n string, in line 9

            • With the Columns++ and Search++ Regex the ^abc regex counts 4 matches. So a beginning of line is seen :

              • At the very beginning of the file

              • After a LF character

              • After a CR character

              • After a \r\n string, in line 9

            • With Search++ ICU, the ^abc counts 9 matches. So a beginning of line is seen, like with N++ and MultiReplace and, also, after a VT character.

            Seemingly, in ICU mode, the \r\n couple is not considered as one char. Thus, the regex ^\n|Q do see the CR character as a beginning of line (^). However, in this case, I do not understand why the ^\n regex finds nothing at all !


            THIRD test

            • Note that in Search++ Settings, I checked the Always unmark all text before Mark command. (Otherwise add to existing marked text.) option.

            • In addition, I chose the Bookmark lines when marking text option, in the Tools menu

            Thus, in all cases, all the marked text and all the bookmarked lines should be wiped out before a new Mark command.

            Using my example text, described above, I marked, in ICU mode, all the lines which match the regex ^\n|Q, as shown in the snapshot below :

            56243346-2c25-4586-b5c0-c44da1de1747-image.jpeg

            Then, I simply searched for the ^\n regex. As noticed above, this regex finds nothing. However, after clicking on the Mark in Whole document option, the previous marks and bookmarks remain unchanged ?

            I expected no marked text and no bookmrked line ! I suppose it’s a bug.

            Best Regards

            guy038

            CoisesC Thomas KnoefelT 2 Replies Last reply Reply Quote 1
            • guy038G Online
              guy038
              last edited by

              Hello, @coises,

              I said, at the end of the second test section :

              Thus, the regex ^\n|Q do see the CR character as a beginning of line (^)

              I didn’t express myself clearly. I should have written :

              Thus, the regex ^\n|Q do see the transition between each CR character to its next LF character as a beginning of line (^)

              BR

              guy038

              1 Reply Last reply Reply Quote 1
              • CoisesC Online
                Coises @guy038
                last edited by Coises

                @guy038 said:

                The only way to bookmark the last empty line 10, that I found, is to use the regex ^((?!test).)*$

                Yes, that’s the bug you found. The precise conditions are that if a match includes the last character in the file, Search++ fails to check for the possibility of a null match at the end of the file. That will be fixed in the next release.

                • Using Search++, in mode Regex, this regex counts 8 matches, whose line 1 with ONLY a small blue triangle at bottom of line 1 and the last empty line 10 with the calltip ^ zero length match

                Yes. Where possible I use the small triangle to mark a zero-length match. There are two situations where that isn’t possible, due to Scintilla limitations: when the match is at the end of a line, before the line ending characters, and line ending characters are not shown; and when the match is at the very end of the file. In those cases, I use the ^ zero length match banner instead.

                I use triangles for zero length matches in the Search++ results list and with the Show command, too. The same limitations apply with Show, but because call tips disappear when you interact with the text, I can’t really do anything about the indicators you can’t see. In the results list you can always see them because I can safely change some settings so that I can use line ending characters that Scintilla considers displayed, but that don’t actually show anything you can see.

                • Using Columns++, this regex counts also 8 matches, but, this time, the first line returns, instead, the calltip ^ zero length match

                Yes, I hadn’t thought of the triangle method yet when I wrote Columns++.

                • Using Search++, in mode ICU this regex counts 7 matches, whose line 1 with ONLY a small blue triangle at bottom of line 1. So the line 10 is not a match

                As best I can tell, ICU’s regex engine does not consider the position following the last character in the document to be the beginning of line, regardless of whether the last character is a line ending character. That makes sense, really, but it doesn’t conform to the way Scintilla lays out lines. So your expression doesn’t match because the circumflex doesn’t match.

                • With N++ and MultiReplace the ^abc regex counts 8 matches. So a beginning of line is seen :

                  • At the very beginning of the file

                  • After a LF character

                  • After a FF character

                  • After a CR character

                  • After a NEL character

                  • After a LS character

                  • After a PS character

                  • After a \r\n string, in line 9

                • With the Columns++ and Search++ Regex the ^abc regex counts 4 matches. So a beginning of line is seen :

                  • At the very beginning of the file

                  • After a LF character

                  • After a CR character

                  • After a \r\n string, in line 9

                • With Search++ ICU, the ^abc counts 9 matches. So a beginning of line is seen, like with N++ and MultiReplace and, also, after a VT character.

                Yes, Columns++ and Search++ Regex do their best (I think they succeed) to treat the same things as line endings that are visible as line endings in Notepad++. So, for example, in your Total_Chars.txt, ^ matches 3 times in Columns++ and in Search++ Regex; it matches 13 times in Search++ ICU; and in Notepad++ native search, if you use Find Next (since Count ignores null matches) and count manually (making sure not to miss the match at the very beginning of the file) there are 7 matches.

                Likewise, (?-s:(?!.))(?s:.) produces 6 matches in Notepad++ (you can use Count for this one), 2 matches in Columns++ and Search++ Regex, and 7 matches in Search++ ICU.

                I’m not sure yet what causes the extra matches for ^ in ICU; it could be something I’ve done wrong in preparing the string, it could be a bug in ICU4C, or could just be something I don’t understand.

                Seemingly, in ICU mode, the \r\n couple is not considered as one char. Thus, the regex ^\n|Q do see the CR character as a beginning of line (^). However, in this case, I do not understand why the ^\n regex finds nothing at all !

                ICU appears to treat ^ differently in different expressions. So far, it looks to me as if, when it is the first thing to match in the only alternative, it treats CRLF as one, but when it’s anywhere else, it treats CR and LF each as line ending characters even when they are together in that order. That’s just from trying to infer a pattern behind what I’ve observed; I can’t find any documentation of such a thing. It could be a bug, or just something I don’t know.

                Then, I simply searched for the ^\n regex. As noticed above, this regex finds nothing. However, after clicking on the Mark in Whole document option, the previous marks and bookmarks remain unchanged ?

                I expected no marked text and no bookmrked line ! I suppose it’s a bug.

                It is working as intended, but perhaps the design is confusing. I had to keep the text in the Settings dialog reasonably brief, but I see I didn’t describe the details correctly in the help.

                Existing marks (and bookmarks, if applicable) are cleared if the command is successful, meaning it finds at least one match. (Internally, it waits until it finds the first match and only clears the existing marks after it has found the first match, but before it marks it.) If there is an error, or if no match is found, nothing is cleared.

                Since you can’t “undo” changes in marks and bookmarks, I thought that at least in the case where someone mistypes a search string resulting in no matches, it would be better not to lose the marks (or selections, or shown text and lines, as the case might be), so one could easily try again.

                1 Reply Last reply Reply Quote 1
                • guy038G Online
                  guy038
                  last edited by guy038

                  Hello, @coises,

                  You said, at the end of my last post :

                  Existing marks (and bookmarks, if applicable) are cleared if the command is successful, meaning it finds at least one match. (Internally, it waits until it finds the first match and only clears the existing marks after it has found the first match, but before it marks it.) If there is an error, or if no match is found, nothing is cleared.

                  I verified that if I, simultaneously, check the third options :

                  • ☑ Always clear selections before Select command. (Otherwise add to existing selections.)

                  • ☑ Always unmark all text before Mark command. (Otherwise add to existing marked text.)

                  • ☑ Always hide all lines and remove Show style from text before Show command

                  Nothing is cleared if current search does not find any match or finds an error, regarding selections, marks, and shown text

                  Therefore, I think it would be a good idea to mention this fact in the Search++ documentation, within the cleaning section of Search++ settings !

                  Best Regards

                  guy038

                  1 Reply Last reply Reply Quote 1
                  • Thomas KnoefelT Offline
                    Thomas Knoefel @guy038
                    last edited by

                    Hi @guy038,

                    MultiReplace doesn’t use a regex engine of its own: its searches run through Notepad++'s Scintilla, i.e. through Notepad++'s Boost regex engine. So ^ and $ see the same line breaks as the native Find dialog, which with Boost also includes FF, NEL, LS and PS. That is intended: MultiReplace should always find the same matches as Notepad++ itself.

                    1 Reply Last reply Reply Quote 1
                    • CoisesC Online
                      Coises
                      last edited by Coises

                      @astewart77, @guy038 and anyone else who cares to comment:

                      As noted, I still didn’t get the Synchronize marks and bookmarks sub-menu right in version 0.7.1. I am interested if you have any preference among these three possibilities:

                      menuimages.png

                      or if you would suggest anything else?

                      As elsewhere, right-click would be equivalent to Shift+click.

                      1 Reply Last reply Reply Quote 0
                      • guy038G Online
                        guy038
                        last edited by guy038

                        Hi, @coises, @astewart77 and All,

                        Ah,…, personally, I’m voting for the Methiod 3 as we are already used to the Shift behaviour. In addition, it remains only three items in the menu which appears more clear !

                        This is coherent with the Mark selected text (Shift: clear first) and Mark shown text (Shift: clear first) options, in the Tools menu !

                        Best Regards,

                        guy038

                        astewart77A 1 Reply Last reply Reply Quote 1
                        • astewart77A Offline
                          astewart77 @guy038
                          last edited by astewart77

                          I also vote for Method 3.

                          Method 2 has too much text changing from use to use. For me, the only advantage to Method 1 is being a bit easier to automate with AutoHotkey.

                          1 Reply Last reply Reply Quote 2

                          Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                          Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                          With your input, this post could be even better 💗

                          Register Login
                          • First post
                            Last post
                          The Community of users of the Notepad++ text editor.
                          Powered by NodeBB | Contributors