4 ms·
Edge cases like this will just get more common as unicode keeps getting more complex. There was a fun slide in this talk[1] that suggests unicode might be turin
by rom-antics 4y ago
Edge cases like this will just get more common as unicode keeps getting more complex. There was a fun slide in this talk[1] that suggests unicode might be turing complete due to its case folding rules.
I miss when Unicode was just a simple list of codepoints. (Get off my lawn)
[1]: https://seriot.ch/resources/talks_papers/20171027_brainfuck_dominos.pdf https://seriot.ch/resources/talks_papers/20171027_brainfuck_...
- GuB-42 4y agoI think we need some kind of standard "Unicode-light" with limitations to allow it to be used on low specs hardware and without weird edge cases like this. A bit like video codecs that have "profiles" which are limitations you can adhere to to avoid overwhelming low end hardware. It wouldn't be "universal", but enough to write in the most commonly used languages, and maybe support a few, single codepoint special characters and emoji.
- themerone 4y agoYou can't express the most commonly used languages without multiple code point graphemes. If you want to eliminate edge cases you would need to introduce new incompatible code points and a 32 or 64 bit fixed length encoding depending on how many languages you want to support.
- GuB-42 4y agoExtra pages for these extra code points wouldn't seem far fetched to me. We already have single code point letters with diacritics like "é", or a huge code page of Hangul, for which each code point is a combination of characters. As for encoding, 32 bits fixed length should be sufficient, I can't believe that we would need billions of symbols, combinations included, in order to write in most common languages, though I may be wrong. Also, "limiting" doesn't not necessarily means "single code point only", but more like "only one diacritic, only from this list, and only for these characters", so that combination fits in a certain size limit (ex: a 32 bit word), and that the engine only has to process a limited number of use cases.
- sbierwagen 4y agoWhy bother with a standard, just deny by default and whitelist the codepoints you care about. Plenty of software already does that-- HN itself doesn't allow emoji, for example.
- rom-antics 4y agoI don't think HN is deny by default. Most emojis are stripped, but some still get through. I don't know why 🆔 would be whitelisted
- zerocrates 4y agoIs it just that whole block that's allowed? I feel like some of the legacy characters that have been kind of "promoted" to emoji I've seen allowed here. Let's see, is the all-important 🆗 allowed?
- GuB-42 4y agoThe point of having a standard is to know which ones to deny. Make a study over a wide range of documents about the characters that are used the most, see if there are alternatives to the characters that don't make it, etc... This is unlike the HN ban on emoji, which I think is more of a political decision than a technical one. Most people on HN use systems that can read emoji just fine, but they decided that the site would be better without it. This would be more technical, a way to balance inclusivity and technical constraints, something between ASCII and full Unicode.
- zokier 4y agoThe subset of codepoints that are included in NFKD make probably a decent starting point for such standard, maybe if you want to be even more restrictive then limit to BMP.
- klausa 4y agoThe _entire point_ of Unicode is to be _universal_, you're just suggesting we go back to pre-Unicode days and use different code-pages. This, for reasons stated in already posted responses, and to put it mildly, does not work in international context (e.g. the vast majority of software written every day).
- WorldMaker 4y agoThese "edge cases" have always existed in Unicode. Languages with ZWJ needs have existed in Unicode since the beginning. That emoji put a spotlight on this for especially English-speaking developers with assumptions that language encodings are "simple", is probably one of the best things about the popularity of emoji.
- Dylan16807 4y agoUnicode always had combining characters. Is this so different from accent marks disappearing? And the Hangul pieces were there from 1.0/1.1.