Hello, @troglo37 and All,
OK. So let’s start off slow !
First of all, I’d like to check with you to make sure we’re on the same page about our goal :
If we begin with this text :
Our friends Sarunas, Dorde, Orjan, Pawel, Stefania, Angel and Ugur had already left
Our friends Šarūnas, Đorđe, Ørjan, Paweł, Ștefania, Ángel and Uğur had already left
Our friends Šarūnas, Đorđe, Ørjan, Paweł, Ștefania, Ángel and Uğur had already left
As you can notice these lines are quite similar and differ only because of some accents in the first names :
The first line does not contain any accent
The second line contains some characters with an accent :
Š = 0160
ū = 016B
Đ = 0110
đ = 0111
Ø = 00D8
ł = 0142
Ș = 0218
Á = 00C1
ğ = 011F
The third line contains — when it was possible to find an equivalent — a classical ASCII letter immediately followed with a combining diacritical mark — an accent — thus two characters for one glyph. However, the letters Đ, đ, Ø and ł, without equivalent, are single characters only, like above !
S + 030C = Š
u + 0304 = ū
Đ = 0110
đ = 0111
Ø = 00D8
ł = 0142
S + 0326 = Ș
A + 0301 = Á
g + 0306 = ğ
Now, I assume you’d be interested in finding any of these three syntaxes — or others like them — at the same time, in your file(s), right ?
@troglo37, just tell me if this is your goal ?
See you later !
Best Regards,
guy038