Community
    • Login

    Search++: A work in progress

    Scheduled Pinned Locked Moved Notepad++ & Plugin Development
    114 Posts 7 Posters 33.9k Views 2 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • guy038G Offline
      guy038
      last edited by guy038

      Hi, @coises and All,

      Sorry, but I came across a serious bug regarding the ICU search, since Search++ version v0.6

      • The Find in Selection, Find in Marked Text and Find in Whole Document options

      • The Find All in Selection, Find All in Marked Text and Find All in Whole Document options

      Triggers the closure of Notepad++, without saving any configuration file and unsaved tabs


      After some quick tests :

      • The Select, Mark and Show options,

      • The Count in Selection, Count in Marked Text and Count in Whole Document

      Seems to work correctly in ICU mode

      BR

      guy038

      CoisesC 1 Reply Last reply Reply Quote 0
      • CoisesC Offline
        Coises @guy038
        last edited by

        @guy038 said:
        I came across a serious bug regarding the ICU search, since Search++ version v0.6

        Thank you for the detailed report. I can reproduce it… in the release version. It doesn’t happen in a debug version (which uses exactly the same code); so it’s a subtle error somewhere. I will find it.

        I changed some existing code when I added Find in Files. You observation that this fault began in 0.6 will help me know where to look.

        1 Reply Last reply Reply Quote 0
        • CoisesC Offline
          Coises @guy038
          last edited by

          @guy038 said:

          • Would it be possible, in the Search++ Results to prevent the displaying of the last character ( an ? sign inside a rectangle ) in all lines ( Search line, Document lines and Matched lines ) ?

          Huh. I know what must be causing that, but I don’t know why it is causing it (on your system, but not on mine). (I set line ending characters to be displayed, but then to display with zero opacity. I do this so the marks I place to indicate a zero-length match will still show when the match occurs at the end of a line. The mark must be placed on the following character, and if the character isn’t shown, neither is the mark.)

          Would you tell me what font you use for Notepad++ documents (the Search++ Results window copies that from the document window) and what you have at Settings | Preferences… | MISC. for rendering mode (GDI or DirectWrite)?

          If I can figure out why you see the characters, hopefully I can figure out what to change so that you (and others with the same configuration) won’t see them.

          Natively, you can swap the keyboard focus between the N++ Search results and the Current document window by hitting the F7 key. I do not see an equivalent with your Search++ plugin

          The problem is that Search++ can’t (or shouldn’t… see below) control what shortcut keys do when focus is in Notepad++.

          From the Search++ Results window, Ctrl+N will take you to the document. To switch from the document to the Results, though, I’d have to add a Search++ Results entry on the Plugins | Search++ menu, and then you’d have to set a shortcut for that in the Notepad++ shortcut mapper. You probably wouldn’t want to give up Ctrl+N, though.

          The only way to get from Notepad++ to Search++ (or any plugin) is through a shortcut that’s either manually added in Shortcut Mapper or added by default in the plugin’s setup. (I’m still not sure how conflicts are resolved when a plugin defines a shortcut key that is already defined by default, by the user or by another plugin. I did define a few shortcut keys in Columns++ for manipulating rectangular selections, but it still bothers me that it feels a little… arrogant… to do that. My plugin might not be as important to the user as something else that also wants those keys.)

          I am open to suggestions on how to manage this. I agree that consistency between navigation keys used when focus is in the plugin and when focus is in Notepad++ would be desirable, I just don’t have a plausible design for it.

          1 Reply Last reply Reply Quote 0
          • guy038G Offline
            guy038
            last edited by guy038

            Hi, @coises,

            Practically, I installed your plugin Search++ on a 8.9_x64 version of N++

            As you can see, in the screenshot below :

            f4e33f66-0743-43a9-9a4f-29e9c401507f-image.jpeg

            • I did not hit the Show All characters option ( ¶ )

            • I did a search on a specific folder ( ^Music )

            • Most of the time, I use the Consolas font ( Consolas / Version 7.00 / Disposition **Open Type** signé numériquement **True Type Contours )

            But you found out the culprit : it’s just because the GDI ( most compatible) option was selected !. Luckily, for any other option of the rendering mode, after closing an re-starting N++, this special character does not display, anymore !


            Regarding the problem about swapping between the Search++ Results and the Current Document window, do not bother about it :-))

            May be, you remenber that I personally used the Ctrl + Shift+ N shortcut to open the Search... option. Thus, each use of the Ctrl + Shift + N, alternatively, opens or closes my Search... docked panel

            So, I was able to deduce a way ( rather long, of course ) to do that :

            • Firstly, use Ctrl + N shortcut to switch to the Current Document window

            • Secondly, use the Ctrl + Shift + N shortcut to put the focus on the Search.. panel

            • Thirdly, use the Ctrl + H shortcut to get the focus on the Search++ Results panel, again !

            Even, on my French keyboard, the H key is quite close to the N key !

            BR

            guy038

            CoisesC 1 Reply Last reply Reply Quote 1
            • CoisesC Offline
              Coises @guy038
              last edited by

              @guy038 said:

              But you found out the culprit : it’s just because the GDI ( most compatible) option was selected !. Luckily, for any other option of the rendering mode, after closing an re-starting N++, this special character does not display, anymore !

              Thanks for documenting that!

              It should appear correctly in GDI mode too, of course. I suspect the problem will turn out to be that GDI mode doesn’t support alpha transparency, so my setting the opacity to zero is doing nothing. I’ll have to find another way.

              1 Reply Last reply Reply Quote 0
              • CoisesC Offline
                Coises
                last edited by

                Search++ version 0.6.2 is available:

                • Fix a critical error that caused ICU searches to crash Notepad++.
                • Fix unwanted characters appearing at the ends of lines in the Results list with Notepad++ rendering mode GDI.
                • Fix regression of line ending symbols being in text color instead of white space color, and switch the CRLF symbol in GDI mode to one that hopefully renders more consistently.

                Thanks to @guy038 for catching the ICU error and the unwanted characters in the results list in GDI mode.

                (The ICU error appears to have been due to an uninitialized variable. The faulty code has been there from the beginning. As chance would have it, up until now it just happened to land on memory that contained an innocuous value. Hopefully I didn’t miss anything else similar…)

                1 Reply Last reply Reply Quote 2
                • guy038G Offline
                  guy038
                  last edited by guy038

                  Hello, @coises and All,

                  Again, thanks for quickly fixing the different issues ( and some others, at the same time ! ). So far, everything looks OK ;-))


                  Now, I did a test, searching in all files of my D: USB key, with :

                  • The native N++ Find in Files dialog

                  • The Search++ Search in Files dialog

                  • The MultiReplace Find All in Files command

                  To be sure to process fair tests, I unplung my USB key, re-start my laptop completely, replug my USB key and starts a new N++ session, before running these three Find in Files operations, in turn !

                  I searched for the Fi string, with the Match case and the Regular expression options checked. I got :

                  • 27,112 matches in 833 files of 1,761 searched for the N++ Search Results panel, in 1m 37s

                  • 27,112 matches in 833 files of 1,761 files for the Search++ Search++ Results, in 2m 21s

                  • 27,087 hits in 835 file(s) for the MultiReplace Search results panel, in 1m 36s

                  However, when you re-run this same search, without exiting N++, right after the first search, we get the job done in about :

                  • 6s for the N++ search

                  • 5.2s for the MultiReplace search

                  • 1,1s for the Search++ search

                  So, apparently, the results of the first search seem to speed up all the next same searches, whatever the search engine used ( native one or plugins ones ) !

                  Best Regards,

                  guy038

                  P.S. :

                  I understood the small differences which occur in the MultiReplace search ( 27,112 - 25 ) matches in ( 833 + 2 ) files. But, I prefer, first, to study more closely the MultiReplace plugin, which is very powerful too and reply later to @thomas-knoefel. It concerns the encoding of 4 files which is not recognized properly and which occurs only when the Find in Files command is used. The Find All and Find in Docs commands, themselves, give correct results for these 4 files !?

                  CoisesC 2 Replies Last reply Reply Quote 1
                  • CoisesC Offline
                    Coises @guy038
                    last edited by

                    @guy038 said:

                    So, apparently, the results of the first search seem to speed up all the next same searches, whatever the search engine used ( native one or plugins ones ) !

                    My guess is that the total size of all the files on your USB key was well less than the amount of memory available to Windows. If so, then (I think) that when you repeated a test, almost nothing was read from the USB key: Windows already had the whole drive cached in memory.

                    It makes sense that in that circumstance, Search++ would be several times faster, because it uses multiple threads for search in files; I’m pretty sure both native Notepad++ search and MultiReplace are strictly single-threaded. So when the actual searching, rather than reading the files, is what takes most of the time, Search++ is using all cores of your CPU and the others are using just one.

                    I do wonder, though, why Search++ is so much slower on the first execution. I suspect that something in the way I’m reading files (trying to get several at once, since multiple threads are running) is reducing throughput from the USB key. There is probably a way to make that better. Whether I can figure out what it is remains to be seen.

                    Thanks for your testing!

                    1 Reply Last reply Reply Quote 1
                    • CoisesC Offline
                      Coises @guy038
                      last edited by Coises

                      @guy038 said:

                      I understood the small differences which occur in the MultiReplace search ( 27,112 - 25 ) matches in ( 833 + 2 ) files. But, I prefer, first, to study more closely the MultiReplace plugin, which is very powerful too and reply later to @thomas-knoefel. It concerns the encoding of 4 files which is not recognized properly and which occurs only when the Find in Files command is used. The Find All and Find in Docs commands, themselves, give correct results for these 4 files !?

                      As best I can make out, like Search++, MultiReplace processes searches against documents open in Notepad++ using the copy of the document that Notepad++ has loaded into Scintilla. (I can’t really think of another way it could be done.) That means that for existing files, Notepad++ has already determined the encoding and either loaded it as ANSI or UTF-8, or translated it to UTF-8. (Notepad++ never sets Scintilla to use any encoding other than UTF-8 or the system default code page; anything else is translated to UTF-8 and translated back on writing.)

                      So it’s no surprise that we’re both using the same encoding as Notepad++ for Find All and Find in Open Documents, because we’re just using the work Notepad++ has already done.

                      I think MultiReplace uses an invisible Scintilla to process searches in files. It could be that it hasn’t entirely replicated the labyrinthine logic Notepad++ uses to determine the encoding of a file (including the effect of Preferences | MISC. | Autodetect character encoding).

                      Search++ does not duplicate that logic either, but it approaches the problem in a different way. Rather than using an invisible Scintilla control, Search++ uses its customized version of Boost::regex directly on the data in the file buffer. So I never translate the file encoding to anything else. I do have to determine the encoding, though, as the choice of iterator for Boost::regex depends on that.

                      There are limitations, not yet formally documented, to the way Search++ handles encodings for Search in Files:

                      1. UTF-16 files without a byte order mark are not recognized as such. (They’ll be misread as something else.)
                      2. UTF-8 files without a byte order mark that contain any invalid UTF-8 sequences will be processed using the system default (ANSI) code page.
                      3. Pure ASCII files are processed as UTF-8. (This doesn’t matter for Find, but I will have to review it when I implement Replace, since someone might include non-ASCII characters in the replacement text. I think there will have to be a user control to determine whether to promote ASCII to ANSI or to UTF-8.)
                      4. Files without a byte order mark that are not pure ASCII and contain any invalid UTF-8 characters are processed using the system default code page. No attempt is made to detect whether a different legacy code page is more likely to be correct, or is explicitly declared within the file.

                      The sad fact about file encoding is that in most cases there is no algorithmic way to be absolutely certain about what it is. (The exceptions are pure ASCII — but then, if any non-ASCII characters are introduced, there is no way to know how those should be encoded — and some file types, like HTML and XML, that can include a declaration of their own encoding, the declaration itself being in plain ASCII. Byte order marks are, in practice, another exception: they could occur in a legacy encoding, but they would be so uncommon at the start of a file that it’s safe to assume they are definitive, or the file has been purposely engineered to cause mis-detection.)

                      I will probably improve handling of file encodings as Search++ evolves. I doubt that I will attempt to duplicate exactly what Notepad++ does.

                      1 Reply Last reply Reply Quote 1
                      • CoisesC Coises referenced this topic
                      • guy038G Offline
                        guy038
                        last edited by

                        Hello, @coises and All,

                        You said :

                        There are limitations, not yet formally documented, to the way Search++ handles encodings for Search in Files:
                        
                        1. UTF-16 files without a byte order mark are not recognized as such. (They’ll be misread as something else.)
                        
                        2. UTF-8 files without a byte order mark that contain any invalid UTF-8 sequences will be processed using the system default (ANSI) code page.
                        
                        3. Pure ASCII files are processed as UTF-8. (This doesn’t matter for Find, but I will have to review it when I implement Replace, since someone might include non-ASCII characters in the replacement text. I think there will have to be a user control to determine whether to promote ASCII to ANSI or to UTF-8.)
                        
                        4. Files without a byte order mark that are not pure ASCII and contain any invalid UTF-8 characters are processed using the system default code page. No attempt is made to detect whether a different legacy code page is more likely to be correct, or is explicitly declared within the file.
                        

                        Thanks for this extra information !

                        FYI, during all my tests, the Autodetect character encoding option, in Settings > Preferences… > MISC, was checked and all the files, described in the P.S. section of my previous post, contain a Byte Order Mark ( BOM ). Here is the list of all UTF-16 files of my D: USB key

                            •------------•--------------------------------------------------------------•-----------------•---------------------•--------------------•
                            |            |                                                              |                 |    Multi-Replace    |   Search++ / N++   |
                            |    BOM     |                             File                             |    Encoding     |     Matches 'Fi'    |    Matches 'Fi'    |
                            •------------•--------------------------------------------------------------•-----------------•---------------------•--------------------•
                            |   FE FF    |  D:\Plane_0_UCS-2_BE.txt                                     |  UTF-16 BE BOM  |          2          |          0         |
                            |   FE FF    |  D:\Planes_0+1_UTF-16.txt                                    |  UTF-16 BE BOM  |          2          |          0         |
                            •------------•--------------------------------------------------------------•-----------------•---------------------•--------------------•
                            |   FF FE    |  D:\862_x64\plugins\NppExec\doc\NppExec\NppExec_HelpAll.txt  |  UTF-16 LE BOM  |          0          |         31         |
                            |   FF FE    |  D:\Plane_0_UCS-2_LE.txt                                     |  UTF-16 LE BOM  |          2          |          0         |
                            |   FF FE    |  D:\862_x64\plugins\Config\npec_cmdhistory.txt               |  UTF-16 LE BOM  |          0          |          0         |
                            |   FF FE    |  D:\862_x64\plugins\Config\npes_last.txt                     |  UTF-16 LE BOM  |          0          |          0         |
                            •------------•--------------------------------------------------------------•-----------------•---------------------•--------------------•
                            |            |                            Total                             |                 |  3 Files / 6 hits   |  1 File / 31 hits  |
                            •------------•--------------------------------------------------------------•-----------------•---------------------•--------------------•
                        

                        Notes :

                        • The correct results are reported in the last column

                        • My USB key also contains 49 UTF-8-BOM files ( EF BB BF ) correctly detected by all the search engines !


                        Now, based on my test, have you conducted any similar tests of your own, and have you found that searching with Search++ also takes longer compared to the native search in N++ ?

                        Of course, I can run more tests myself, but I imagine your tests would be more valuable and might, perhaps, give you some ideas on how to increase the overall search speed with Search++

                        I’m really curious to hear your own observations !

                        Best Regards,

                        guy038

                        1 Reply Last reply Reply Quote 0

                        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                        With your input, this post could be even better 💗

                        Register Login
                        • First post
                          Last post
                        The Community of users of the Notepad++ text editor.
                        Powered by NodeBB | Contributors