4 ms·
Having a 1:1 correspondence between characters and code points would have been nice.
by seri4l 4y ago
Having a 1:1 correspondence between characters and code points would have been nice.
- samtho 4y agoI think we should have at least an agreed upon subset of Unicode that squashed characters which are deemed too visually similar into singular code points, eliminating the concept of caps vs lowercase (if the language supports it), and remove any lexical control characters to get an analogue if the the “pure” hostname-style charset we enjoy in ASCII that we find in RFCs standards, for example. I’m probably naïve about this, but this is how we treat similar problems within specific applications of the ASCII charset.
- Avamander 4y agoThere certainly are some things that could've done better and I guess still some things we could do better avoiding past mistakes. But I have a strong bad feeling what's here already is going to remain here for the foreseaable future and the best we're gonna get is better handling rather than large changes to unicode itself.
- bawolff 4y agoWell we certainly could do better, we're never be perfect on this. After all, paypa1.com is the same attack and doesn't even use unicode.
- sterlind 4y agothere's valid reasons to have different code points with the same character though. like Russian с and Latin c look identical to me, but fonts might have ligatures or kerning rules that apply to one but not the other.
- bawolff 4y agoDepending on how far you want to take this - 1 and l look the same on many fonts.
- mananaysiempre 4y agoI’m assuming you mean a 1:1 correspondence between code points and glyph shapes. In that case, you won’t have roundtrip conversions with any legacy encoding beyond Latin. For example, every legacy Cyrillic encoding treats the Latin A and the Cyrillic А as different letters. (Lest you try some sort of context-sensitive transform, both are in use as single-letter word: a French verb form and a Russian conjunction, respectively.) Even if we could deal with that, the Cyrillic letter that is written as д (pronounced [d]) when used in printed Russian is written exactly like a Latin single-storey g when used in Bulgarian (never as a double-storey one). This is a problem for your idea even in vacuum, but every legacy encoding also treats this as a mere font difference. As a further example, a and ɑ denote different sounds in the IPA (present simultaneously in French), yet there are fonts where the former looks like the latter (this is frequent in italics, but e.g. regular Futura and Andika do that as well). Those fonts are unsuitable for IPA, of course, but that shouldn’t mean a universal encoding must be unsuitable for it as well. Then there’s the Greek α, which is definitely not an a but kind of like an ɑ. Do you want to distinguish the German Eszett ß (sometimes has a small protrusion on the left due to its origin as a ligature of long S + S/Z) and the Greek β (never has one)? The Greek τ, the Cyrillic т, the Latin m (used as the standard shape for т in Bulgarian), and the Latin(!) m-overbar (used as the standard shape for т in Serbian)? I guess what I’m getting at is that the equivalence classes under “some language’s writing tradition has a glyph for C1 that is very similar to some other language’s writing tradition glyph for C2” end up much larger than anyone would ever want for general text-on-computers use.
- cryptonector 4y agoRound tripping through conversions to other codesets was always only temporarily needed, and for most non-Unicode codesets it was always limited to one or a handful of scripts (e.g., you could not round-trip text that included CJK and Cyrillic through any non-Unicode codeset). So I'm not sure that having that feature was all that important. However, early on it was indeed helpful to obtaining adoption to be able to convert with simple lookup tables.
- deleted 4y ago[deleted]