4 ms·
The second email address uses the Cyrillic small letter O, which renders the same as the Latin small O in almost all fonts. You can see this in your hex dump, t
by FreeFull 6y ago
The second email address uses the Cyrillic small letter O, which renders the same as the Latin small O in almost all fonts. You can see this in your hex dump, too: Instead of a single o, your hex dump shows two bytes for the second character in the second email address.
- inetknght 6y agoYup! I was surprised they're rendered the same. The hex dump didn't lie even though everything until then did.
- hombre_fatal 6y agoThey aren't "lying", it's the same glyph in many fonts. That you think homoglyphs would somehow look different in any of those steps of your post is peculiar to me. l and I are the same in some fonts and they are in the ascii set.
- microcolonel 6y agoI think part of the mistake was separating them to begin with. Far as I can tell, they are literally the same character, with the same origin.
- TheDong 6y agoUnicode, in some cases, did avoid separating characters out, and that was also a mistake that damaged existing alphabets irreparably. See the Han unification [0] effort. There are characters in Japanese and Chinese which are similar, but written differently in each... And they ended up using the same codepoint in unicode for them and relying on different fonts. So now I can't easily quote a japanese sentence in a chinese book without having to use two different fonts, which seems quite silly. Worse yet, there were many common glyphs in use (especially for names) that were unified out of existence. There are literally people who couldn't type their names in unicode. There are a lot of works of text that can't be faithfully OCRd due to not having certain character variants that were unified away. Okay, so why wasn't the cyrillic alphabet unified with the latin one even if the japanese and chinese ones were? Clearly Han unification did much more damage than cyrillic unification would have. Well, the answer is sorta politics. ISO 8859 is what came before Unicode, as far as the unicode consortium is concerned. Since ISO 8859 encoded latin and cyrillic separately, that got carried over for "compatibility". Because the unicode consortium and ISO 8859 were both more western-centric, and CJK users had already dealt with things in a way that was standardized only over there, not in any western ISO standard, of course the unicode consortium would honor the existing ISO standard and ignore the CJK standards. [0]: https://en.wikipedia.org/wiki/Han_unification https://en.wikipedia.org/wiki/Han_unification
- microcolonel 6y agoIdunno, the Han unification wasn't that aggressive. What would really be destructive would be something like 瞭 and 了 sharing a codepoint. Which glyphs were “unified out of existence”? Isn't it more just a matter of using the appropriate typeface or a variation selector? In my limited experience, I've noticed some things being non-unified that I would expect to be unified, like 步 and 歩 (e.g. in 散步 [zh-TW] vs 散歩 [ja-JP]); far as I can tell there isn't any semantic difference between these characters, 新字體 just added a stroke to make one of the radicals more consistent; I guess the idea was that Japanese users mix 新字體 and 舊字體 in text, and JIS character sets had separate codepoints for each from the beginning. As somebody who is learning both Taiwanese Mandarin and Japanese, I don't think it's common for unified characters to differ enough that they would be unrecognizable. I think that no matter how the Unicode Consortium approached this, they would need to set some limits on what gets a separate codepoint; the question is not whether or not to do the Han Unification, the question is when a glyph is actually its own character. As for selecting specific glyphs of a character, if you have a font that even has the variant you're looking for, there is https://en.wikipedia.org/wiki/Variation_Selectors_Supplement https://en.wikipedia.org/wiki/Variation_Selectors_Supplement
- powersnail 6y ago步 and 歩 are similar for sure, and for people aware of what's in Han unification, it's probably not that bad. But the problem is, there are many characters very similar to each other, with a difference of one stroke already. 今 and 令 for example, are both simplified Chinese, but with completely unrelated meaning. So, when you see a character that looks familiar, how can you tell whether you are looking at a character that you don't know or a character you know but rendered in a different language?
- moron4hire 6y agoHow do I look at the sequence of Roman characters “gift” and decide if it is an English-language thing I want vs a German-language thing I don’t? You can’t code your way out of a context trap.
- 6y ago
- inetknght 6y ago> They aren't "lying", it's the same glyph in many fonts. While it might represent the same glyph, it certainly isn't the same sequence of bytes. I think the real failure is that's not made clear.
- hombre_fatal 6y agoThat's what I mean, looking at the bytes is the only way to know. How would you solve homograph/glyph attacks though? One idea is yet another encoding where there are no homoglyphs, only whitelisted diacritic sequences, and there aren't more than one way to assemble the same character ("ó" vs "o"+"´". So tough potatoes for Cyrillic "o", it's forced to use its nearest equivalent: 0x6f Latin "o" in the ascii set.
- inetknght 6y ago> How would you solve homograph/glyph attacks though? One idea is yet another encoding where there are no homoglyphs First, modern operating systems (should?) already provide APIs to canonicalize UTF. Second, perhaps an additional API needs to be created which suggests similarities between characters intended for use by an intelligence (artificial or otherwise...).