Community
    • Login
    1. Home
    2. Popular
    Log in to post
    • All Time
    • Day
    • Week
    • Month
    • All Topics
    • New Topics
    • Watched Topics
    • Unreplied Topics

    • All categories
    • donhoD

      Notepad++ release 8.9.8.1

      Watching Ignoring Scheduled Pinned Locked Moved Announcements
      27
      4 Votes
      27 Posts
      4k Views
      CoisesC
      @Xuân-Thơ-HOÀNG said: @Coises Notepad++ v8.9.8.1 appears to misdetect a valid UTF-8 file without BOM as OEM 855. The file contains valid UTF-8 multibyte characters and can be decoded successfully as UTF-8. When “Autodetect character encoding” is disabled, closing and reopening Notepad++ correctly identifies the file as UTF-8. Re-enabling autodetection causes the same file to be detected as OEM 855 again. For example, string UTF8 Ok but incorrect auto-detection: 🧠 🧠 🧠 🧠 🧠 🧠 🧠 🧠 🧠 🧠 Just my personal opinion, not any sort of “official Notepad++ position”: Autodetect character encoding is an accident waiting to happen. There is no way to consistently and accurately detect character encoding on Windows, because Microsoft, in their infinite wisdom, did not choose to store the character encoding of a file in any sort of metadata. If you really have many files with different legacy character encodings, you will have to intervene manually to get correct results. There is no way around it. Autodetect character encoding might or might not reduce the number of cases where you must manually select the correct encoding, but sometimes it will be wrong. You can use the Encoding | Character sets menu to select the correct encoding before you modify the document in any way. More common is to have just two potentially ambiguous cases: so-called “ANSI,” meaning the system default code page, and UTF-8 with no byte order mark. Outside of contrived examples and files that were supposed to be UTF-8 but contain invalid byte sequences, it is easy to tell those two apart with a very high rate of accuracy. (However, Notepad++ has a bug that can cause UTF-8 files to be mis-detected as ANSI.) That works better if you do not check Autodetect character encoding.
    • Dean-CorsoD

      How to normalize fancy Unicode text back to regular text?

      Watching Ignoring Scheduled Pinned Locked Moved Help wanted · · · – – – · · ·
      29
      0 Votes
      29 Posts
      14k Views
      CoisesC
      @Screen-White said: plugin For what it’s worth, I have a plugin called Unicode Normalize that can normalize selected text to any of the four standard Unicode normalization forms. There’s a short discussion of it here. This plugin uses Windows’ NormalizeString function. It could probably be improved by using ICU4C. I haven’t done a lot of work on it; it’s just something I threw together to solve a problem. The search problem is one I hope to solve someday in my work-in-progress called Search++, but it will be a while before I get to that.