4 ms·
The author thinks that by "fooling" search engines and llms, using obfuscated codepoints is "Unicode's" revenge! Unicode is a truly gift to humanity, as it has
by yannis 3y ago
The author thinks that by "fooling" search engines and llms, using obfuscated codepoints is "Unicode's" revenge! Unicode is a truly gift to humanity, as it has so far encoded the majority of the worlds scripts, ancient or modern, brought order from the chaos of earlier encodings. The author forgets that humans write to communicate with other humans. If search engines are confused by such antics let it be.
- avgcorrection 3y agoExactly. And to go back to Hungarians:[1] the alternative to having one large standard for all text is to have multiple encodings for different languages and peoples. And that’s what we did for decades before Unicode which was apparently a mess. And of course there certainly are Hungarians that also write in Thai. So what about mixed Hungarian/Thai text? Should that be a multiple (interleaved) encoding text? I mean as an alternative to “149,813 ‘characters’” that “any one human is [not] likely to use [most of]”. [1] https://news.ycombinator.com/item?id=39066135 https://news.ycombinator.com/item?id=39066135
- yannis 3y agoHere is an example of Hungarian and Thai https://hu.wikipedia.org/wiki/Thai_%C3%ADr%C3%A1s https://hu.wikipedia.org/wiki/Thai_%C3%ADr%C3%A1s
- linguae 3y agoI remember this for the Japanese language back in the early 2000s when I started learning the language. Nowadays Japanese websites mostly use UTF-8, but back in 2000 there was a mixture of JIS, Shift-JIS, and EUC content, and sometimes browsers had difficulties automatically determining the encoding, resulting in rendering “mojibake” until I had to manually specify the encoding. Unicode is a great thing and has made computing more accessible to those who read and write in languages that don’t use the Latin alphabet.