17 ms·
Unicode Is Awesome
- jagracey 7y agoUnicode reverse character: 'hello \u{202e} world'; 'hello dlrow' // Visual equivalent
- jagracey 7y agoThe Emoji emoji modifiers are pretty cool. - Skin color modifiers - Character combiners: - man [ZWJ] woman [ZWJ] boy [ZWJ] girl === family of 4
- johannes1234321 7y ago"How to shrink a family using MySQL" or: "Unicode Emojis, Code Points, and Grepheme Clusters" https://twitter.com/johannescode/status/1183716981612208128?s=19 https://twitter.com/johannescode/status/1183716981612208128?...
- FisDugthop 7y agoPage doesn't render without JS enabled. Enabling JS causes questionable CSP requests.
- euske 7y agoI mean, awesome for whom? It might be awesome for end users as you can type in or copy/paste things without caring which language you're using. But for programmers, Unicode is a bloated monstrosity and a source of endless nightmare. Eventually, it's not going to be awesome for end users either because it will be plagued by a lot of (subtle) inconsistencies. Unicode looks a lot like a leaky abstraction to me (because of poor foresight), and it's getting worse each year.
- jcranmer 7y agoIf you think Unicode is a "bloated monstrosity and a source of endless nightmare," what would you remove from Unicode? And if you're going to respond "emoji", I'll point out that removing emoji doesn't actually remove anything that makes text processing with Unicode difficult, just makes it more likely that people will assume that what works for English works for everybody. (Side note: it is not possible to accurately represent modern English text solely with ASCII, as English does contain several words with accented characters, such as façade and résumé).
- mr_toad 7y agoMost resumes I’ve seen don’t even bother with the accents. Most are written in Word on Windows, and I’d guess that most people don’t even know how to access the accented characters.
- jancsika 7y ago> such as façade and résumé That's simple: just url encode. Compare: www.façebook.com to www.fa%C3%A7ebook.com The second one is way easier to comprehend than the first.
- bawolff 7y agoYou mean www.xn--faebook-vxa.com of course :P
- userbinator 7y agoIf you extend ASCII to CP1252, which is the most common encoding besides/before UTF-8 became common, then you do get those accented characters (and that's likely responsible for the popularity of '1252.) In fact, the first 256 characters of Unicode are almost identical to CP1252. I'm pretty sure that's not a coincidence.
- pdonis 7y ago> the first 256 characters of Unicode are almost identical to CP1252. I'm pretty sure that's not a coincidence. That depends on whether you consider the fact that Windows CP 1252 is almost identical to Latin-1 (ISO-8859-1), which is exactly the first 256 characters of Unicode, to be a coincidence.
- jrochkind1 7y agoUnicode is pretty amazing. People REALLY like to complain about unicode, but where it's complicated, it's because the _problem space_ is complicated. Which it is. People are actually complaining that they wish handling global text wasn't so complicated, like, that humans had been a lot simpler and more limited with their inventions of alphabets and how they were used in typesetting and printing and what not, and that legacy digital text encodings historically had happened differently than they did, they're not actually complaining about unicode at all, which had to deal with the hand of cards it was dealt. That unicode worked out as nice a solution as it is to storing global text is pretty amazing, there were some really smart and competent people working on it. When you dig into the details, you will be continually amazed how nice it is. And how well-documented. One real testament to this is how _well adopted_ Unicode is. There is no actual guarantee that just because you make a standard anyone will use it. Nobody forced anyone to move from whatever they did to Unicode. (and in fact most eg internet standards don't force Unicode and are technically agnostic as to text encoding). That it has become so universal is because it was so well-designed, it solved real problems, with a feasible migration path for developers that had a cost justified by it's benefits. (When people complain about aspects of UTF-8 required by it's multi-facetted compatibility with ascii, they are missing that this is what led to unicode actually winning). The OP, despite the title, doesn't actually serve as a great argument/explanation for how Unicode is awesome. But I'd read some of the Unicode "annex" docs -- they are also great docs!
- kazinator 7y agoNo, where Unicode is complicated is where the Unicode people decided to make it complicated to bolster their egos, to the detriment of everyone downstream of them. Like with most standardization, the people at the helm are the wrong people with the wrong motivations.
- ori_b 7y agoInteresting statement. Other than maybe han unification, what would you do differently?
- 7y ago
- flohofwoe 7y agoIMHO the article should mention that UTF-16 was (more or less) a hack to fix Windows and some other systems which didn't see the light and use UTF-8 from the start. UTF-16 has all the disadvantes of UTF-8 (variable length) and UTF-32 (endianess), but none of the advantages (encoding as endian-agnostic, 7-bit ASCII compatible byte stream like UTF-8, or a fixed-width encoding like UTF-32). UTF-16 should really be considered a hack to talk to (mainly) Windows APIs. Also, obligatory link to: https://utf8everywhere.org/ https://utf8everywhere.org/
- dfox 7y agoThe main point that should be emphasised is that any encoding with fixed size unicode codepoints is mostly unnecessary as you mostly don’t care about the codepoints but about how the resulting glyphs or even glyph runs look like. My experience is that if you want to implement efficient unicode-aware text editor then the right datastructure is list of lines and you have to simply forget about gap buffers, ropes and what not (unless you really care about 32k+ lines/paragraphs, which is when rope-style representation starts to make sense as long as the breaks match unicode semantics)
- wvenable 7y agoWindows and many other operating systems and languages (Java) got on board with Unicode back when the character set would fit in 16bits. The character set originally used was UCS-2 (not UTF-16). UTF-16 came next to extend the Unicode character set beyond 65536 code points. UTF-8 wasn't even invented until well after all these operating systems and languages deployed Unicode. They didn't see the light of day to use UTF-8 because they didn't have a time machine to make that possible.
- flohofwoe 7y agoI actually checked a while ago when UTF-8 was created, and it was just around the same time when Windows NT was developed with 16-bit "early" Unicode support. UTF-8 was created in September 1992 [1], and Windows NT came out mid 1993, but I guess it was too late for Windows to change to UTF-8 (and I guess the advantages of UTF-8 haven't been as clear back then). But IMHO there's no excuse to not use UTF-8 after around 1995 ;) [1] https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt
- akdor1154 7y agoUnicode is great, emoji are a (technically impressive) monstrosity.
- WorldMaker 7y agoEmoji are an almost critical Unicode democratizing need. They aren't doing anything that other languages encoded with Unicode don't already do (and haven't already done since the beginning of Unicode). This article itself points out several key existing relatives, such as how Arabic, one of the most common and important written languages in the world, or the very important CJK family of written languages, used ZWJ and ZWNJ well before Emoji made it "cool" to other parts of the world, most especially the English-writing contingent that has long thought of Unicode as simply "ASCII plus a bunch of other stuff I might never use". Suddenly a lot of English documents have embedded emoji that deeply matters to the writers, and there are fewer excuses to treat Unicode as "ASCII+" and more cases where doing so is not only wrong (broken surrogate pairs, incorrect codepoint analysis for ZWJ, etc), but very visibly wrong in a way that users care and complain about it.
- johncolanduoni 7y agoExcepting some of the weird combining mark tricks used for flags and some more straightforward modifiers for skin tone, they’re basically a bunch of inert codepoints that you can get away with just popping in a high-res PNG for. No complicated typesetting, not even any kerning. If you handle surrogate pairs correctly (which is also needed for some real, widely used languages) common emojis will work fine. I don’t see what’s technically impressive about them, care to elaborate?
- RcouF1uZ4gsC 7y ago> Unicode is simply a 16-bit code - Some people are under the misconception that Unicode is simply a 16-bit code where each character takes 16 bits and therefore there are 65,536 possible characters. This is not, actually, correct. It is the single most common myth about Unicode, so if you thought that, don't feel bad. Verity Stob has a great column https://www.theregister.co.uk/2013/10/04/verity_stob_unicode/ https://www.theregister.co.uk/2013/10/04/verity_stob_unicode... where she says that it is wrong to call this a myth, since that was how it was originally designed. It is better characterized as being obsolete, rather than a myth.
- msla 7y agoIt's a myth about the current version of Unicode. Whether it's true about some obsolete version is hair-splitting at this point.
- deleted 7y ago[deleted]
- deathanatos 7y agoI swear there should be some rule or law about how Unicode articles will inevitably muddle code units/points / grapheme clusters / bytes together. > String length is typically determined by counting codepoints. > This means that surrogate pairs would count as two characters. If you were counting code points, a surrogate pair would be 1. If it's two, you're counting code units. > Combining multiple diacritics may be stacked over the same character. a + ̈ == ̈a, increasing length, while only producing a single character. Not if you're counting code points or code units, which would both produce an answer of "2", and that's a great example of why you shouldn't count with either. The dark blue on black in tables is next to invisible. And then to put that on white on the alternate rows is just eyeball murder. > Since there are over 1.1 million UTF-8 glphys (sic) UTF-8 glyphs twitch; aside from that, I'm really curious how they got that number. In some ways, a font has it easy; my understanding is that modern font formats can do one glyph for acute accent, one glyph for all the vowels/letters, and then compose the glyphs into arrangements for having them combined. (IDK if those are also "glyphs" to the font or not.) But it's less drawing, at least. OTOH, some characters have >1 appearance/"image", AIUI.
- derefr 7y ago> If you were counting code points, a surrogate pair would be 1. If it's two, you're counting code units. And to be explicit as to why that is: surrogate pairs are a feature of the UTF-16 encoding, where two 16-bit code units ("code units" being the lexemes of the decoder) decode to a single Unicode codepoint. I feel like everything to do with Unicode is clearer if you never bring up how it's encoded; or, alternately, if you pretend for the sake of your tutorial that everybody uses UTF-32, so you can just talk about flinging single-code-unit codepoints around as machine-words, the same way ASCII flings single-code-unit codepoints around as bytes. This being basically what Unicode text-handling libraries are doing underneath anyway. After all, from the perspective of the Unicode standard itself, all the stuff below the abstraction of "a codepoint" is implementation detail. The standard has to let the abstraction leak in a few places, like surrogate pairs or BOMs, but these leaks aren't what the Unicode standard is supposed to be "about", and should really be thought of as features of the encodings that have found their way up a layer, rather than features of Unicode per se. Heck, even the categorization of codepoint-ranges into "planes" is just a pragma of UTF-16. Putting these pragma-features front-and-center in a discussion of "what Unicode is", is IMHO entirely backwards.
- akersten 7y agoUnicode is an inspirational standard. We started with so many different character encodings and wound up pretty universally using Unicode. I wouldn't be surprised to see browsers start to drop support for other encodings - who even uses them at this point? Are there any scenarios where you wouldn't use Unicode, other than an every-byte-matters embedded system?
- flohofwoe 7y agoSorry for nitpicking but: Unicode is not an encoding, just (basically) a central registry for numbers, and you can have Unicode strings made of bytes (the UTF-8 encoding), in fact that's the most useful encoding for exchanging text data :)
- TheDong 7y agoThere's a not-insignificant number of Japanese websites that can only correctly display using EUC-JP or ShiftJIS. It seems very Latin/ASCII centric to push for disabling non-UTF-8 encodings, especially since the only reason UTF-8 works so well on ASCII websites is due to its backwards compatibility. If it were the reverse, and UTF-8 were backwards compatible with EUC-JP/CJK, but not ASCII, I doubt you'd be pushing for eschewing other formats since it would break so many english websites.
- jcranmer 7y agoThere is no character in EUC-JP or Shift-JIS that is not in Unicode--the explicit goal of Unicode in its original formulation was to be able to losslessly round-trip any other charset through Unicode, and the initial version of Unicode incorporated the source kanji lists for the EUC-JP/Shift-JIS charsets in their entirety.
- TheDong 7y agoThat's true, but you misunderstood what I meant. The parent comment seemed to be implying that we should drop support for non-utf8 charsets. To me, that rings like saying a website with 'charset=EUC-JP' (such as http://www.os2.jp/ http://www.os2.jp/) should be broken, as in browsers should error out or display a large quantity of black boxes due to it using a non-utf-8 encoding. I'm claiming the only reason the author thinks that's really viable is because in our western-centric world, we see mostly ascii and utf8. Things that, if you flip to only utf-8, both still look fine. CJK websites, on the other hand, that are using the equivalent of ASCII will have to be manually upgraded to display correctly if browsers drop their support. Sure, all their characters can be represented in utf-8, but there's large swathes of websites that will never be updated to a new charset, and it's only a western-centric view that can so blithely suggest breaking them all.
- ken 7y ago> data to be transmitted in a byte, word or double word oriented format (i.e. in 8, 16 or 32-bits per code unit) I don't think I've heard "word" mean "16 bits" since the 1980's ... and apparently neither has Wikipedia: https://en.wikipedia.org/wiki/Word_(computer_architecture)#Table_of_word_sizes https://en.wikipedia.org/wiki/Word_(computer_architecture)#T...
- jontro 7y agoIn Windows world DWORD is 32 bits large: https://docs.microsoft.com/en-us/openspecs/windows_protocols/ms-dtyp/262627d8-3418-4627-9218-4ffe110850b2 https://docs.microsoft.com/en-us/openspecs/windows_protocols...
- pjtr 7y agoAnd WORD 16 bits: https://docs.microsoft.com/en-us/openspecs/windows_protocols/ms-dtyp/f8573df3-a44a-4a50-b070-ac4c3aa78e3c https://docs.microsoft.com/en-us/openspecs/windows_protocols... (QWORD 64 bits)
- WalterGR 7y agoThe concept of the word size of an architecture is different than "word" which has long been used colloquially in computing to mean two bytes. 2 nibbles are a byte. 2 bytes are a word. Edit: "Word" as two bytes may actually be a microcomputer-specific colloquialism.
- grantmnz 7y agoIf you'd like to explore Unicode characters, you can use the Unicode Character finder, a web app I built some years ago: https://www.mclean.net.nz/ucf/ https://www.mclean.net.nz/ucf/ The app allows you to paste in a character to find out more about it, or to search the database of character descriptions to find what you're after. You can link to a specific character to share with your friends and family: https://www.mclean.net.nz/ucf/?c=U+130BA https://www.mclean.net.nz/ucf/?c=U+130BA
- jagracey 7y agoFor exploration, additionally I'd recommend http://shapecatcher.com/ http://shapecatcher.com/ It allows you to draw the shape you are looking for, and with some form of ML, sorts by similarity. It has come in handy a few times for finding the characters I'm unable to describe.
- squaresmile 7y agoUnicode definitely has flaws but that doesn't mean we should throw the baby out with the bathwater and go back to "ASCII and other character sets." There's a reason we moved on from that world. However, I bet we will see another encoding coming up eventually (within 30 years) which solves the problems Unicode currently has and introduces a new set of problems as well. I saw this comment [0] about how that encoding should get started. > Greek, for example, has a lot of special-casing in Unicode. Korean is devilishly hard to render correctly the way Unicode handles it. And once you get into the right-to-left scripts, scripts that sort-of-sometimes omit vowels, or Devanagari (the script used to write a bunch of widely-spoken languages in India), you start needing very different capabilities than what's involved in Western European writing. _The better approach probably would have been to start with those, and work back to the European scripts_ [0] https://www.reddit.com/r/programming/comments/b09c0j/when_zo%C3%AB_zo%C3%AB_or_why_you_need_to_normalize_unicode/eiekg1x/ https://www.reddit.com/r/programming/comments/b09c0j/when_zo... Funnily enough, URLs still can't do actual Unicode.
- BurningFrog 7y agoUnicode URL has serious security problems. The canonical example is google.com vs gооgle.com.
- riquito 7y agoThat was solved years ago by IDN/Punycode (implemented by any browser worth their salt).
- jagracey 7y agoI agree. I'll be releasing an article about this tomorrow. There are in-fact many security ramifications that have not been solved in practice.
- jagracey 7y agoCommented above, but to follow up from yesterday, here is the next post. "Hacking GitHub with Unicode" https://news.ycombinator.com/item?id=21693550 https://news.ycombinator.com/item?id=21693550
- kevmoo1 7y agohttps://youtu.be/dMnPM6z6z40 https://youtu.be/dMnPM6z6z40
- jakeogh 7y agoWhat's the code point for uppercase superscript Z?
- kps 7y agoThere isn't one. Unicode considers superscripting a matter of presentation, which Unicode doesn't cover, except when it does.
- tialaramex 7y agoMore particularly: Presentation variant is not a justification for inclusion in Unicode BUT prior encoding in another character set is. Unicode sets a high priority on roundtripping. The idea is that if you take some data in any one character set X and convert it to Unicode, you should preserve all the meaning by doing this, such that you could losslessly convert it back to encoding X. It's like the wordprocessor problem where users say they only want 10% of the features of a popular wordprocessor but it turns out each user wants a different 10% and so the only way to deliver what they all want is to deliver 100% of the features. Likewise, Unicode has all the weird features of every legacy character set which was embraced BUT it doesn't arbitrarily add new weird features, although you could argue that some of the work done for Unicode has that effect e.g. the way flags work or the Fitzpatrick modifiers. If Unicode had insisted upon never encoding anything that might be a presentation feature, it'd be a long forgotten academic project that never went anywhere and we'd all be using some (probably Microsoft designed) 16-bit ASCII superset today.
- jakeogh 7y agoIs there a realitvely easy way to find the character set that was included for uppercase superscript W? (ᵂ)
- tialaramex 7y agoIn the case of U+1D42 Modifier Letter Capital W I was wrong about the cause, it was in fact specifically added on the rationale that for this purpose (phonetics) the presentation was semantic in nature, and so the plain text (thus Unicode) needed to preserve these symbols which could otherwise be handled by a presentation layer. U+1D42 Modifier Letter Capital W was added in Unicode 4.0 as part of the Phonetic Extensions and Wikipedia provides a long list of Unicode committee paperwork regarding this: https://en.wikipedia.org/wiki/Phonetic_Extensions https://en.wikipedia.org/wiki/Phonetic_Extensions You can see that initially it would have been numbered differently and then over the course of several drafts the proposal evolved until it was assigned U+1D42
- lifthrasiir 7y agoYears ago I've posted support material [1] for Hangul filler mentioned in the article, reproduced below: --- U+3164 HANGUL FILLER is one of the stupidest choices made by character sets. Hangul is noted for its algorithmic construction and Hangul charsets should ideally be following that. Unfortunately, the predominant method for multibyte encoding was ISO 2022 and EUC and both required a rather small repetoire of 94 × 94 = 8,836 characters [0] which are much less than required 19 × 21 × 28 = 11,172 modern syllables. The initial KS X 1001 charset, therefore, only contained 2,350 frequent syllables (plus 4,888 Chinese characters with some duplicates, themselves becoming another Unicode headache). Notwithstanding the fact that remaining syllables are not supported, this resulted in a significant complexity burden for every Hangul-supporting software and there were confusion and contention between KS X 1001 and less interoperable "compositional" (johab) encodings before Unicode. The standardization committee has later acknowledged the charset's shortcoming, but only by adding four-letter (thus eight-byte) ad-hoc combinations for all remaining syllables! The Hangul filler is a designator for such combinations, e.g. `<fliler>ㄱㅏ<filler>` denotes `가` and `<filler>ㅂㅞㄺ` denotes `뷁` (not in KS X 1001 per se). Hangul filler was too late in the scene that it had virtually no support from software industry. Sadly, the filler was there and Unicode had to accept it; technically it can be used to designate a letter (even though Unicode does not support the combinations) so the filler itself should be considered as a letter as well. What, the, hell. [0] It is technically possible to use 94 × 94 × 94 = 830,584 characters with three-byte encoding, but as far as I know there is no known example of such charset designed (thus no real support too). --- I should also mention that early Mozilla (and thus Firefox) had once supported ad-hoc combinations for KS X 1001, got interoperability problems and dropped the support later. Nowadays we treat KS X 1001 as an alias of Windows code page 949 for the sake of compatibility [2]. [1] https://github.com/Wisdom/Awesome-Unicode/issues/4 https://github.com/Wisdom/Awesome-Unicode/issues/4 [2] https://encoding.spec.whatwg.org/#index-euc-kr https://encoding.spec.whatwg.org/#index-euc-kr
- dvfjsdhgfv 7y ago> it's benefits https://news.ycombinator.com/item?id=21679440 https://news.ycombinator.com/item?id=21679440
- nabla9 7y agoUnicode has two really great features. * It names and defines things and sets standard. This seems trivial but is incredibly useful. * Unicode encodings, mainly UTF-8 are good storage format for text (as a data structure for editing text, not so much if you want to be universal). Unicode has one really horrible failing. The 'user-perceived character' (Unicode terminology) is arguably the most important unit in text. Unicode approximates user-perceived characters using set of general rules to define grapheme clusters. A Grapheme cluster is a sequence of adjacent code points that should be treated as a unit by applications. Unfortunately the ruleset and definition is inadequate. Sometimes you need two grapheme clusters to define one unit. If you get UTF-8 encoded and normalized string from somewhere from some unspecified time and era, don't know what application wrote it, using what version of UNICODE standard and what was the locale, you may lose some information. Unicode should have added explicit encoding for user-perceived character boundaries (either fixed grapheme cluster eoncoding or completely different encoding). Let the writing software define it explicitly. It would have been future-proof (new software in the future can understand old strings) and past-proof (ancient software can understand and edit strings written in the future).
- jagracey 7y agoJust posted another Unicode article to HN. "Hacking GitHub with Unicode" https://news.ycombinator.com/item?id=21693550 https://news.ycombinator.com/item?id=21693550