39 ms·
What every software developer must know about Unicode in 2023
- wyldfire 3y ago> “I know, I’ll use a library to do strlen()!” — nobody, ever. The standard library provided by languages like C, C++ is a library. Features like character strings are present and it's a totally reasonable expectation for the length to give you the cluster count.
- AnimalMuppet 3y agoNo, for C and C++, which are close to the hardware, it's totally reasonable to expect strlen() to give you the byte count. You don't allocate memory for buffers based on the cluster count. If you want cluster count, call a different function.
- macintux 3y agoGiven that strlen() predates Unicode by...30 years(?) - it's not terribly surprising that isn't a viable approach.
- bumbledraven 3y ago> what to you think "ẇ͓̞͒͟͡ǫ̠̠̉̏͠͡ͅr̬̺͚̍͛̔͒͢d̠͎̗̳͇͆̋̊͂͐".length should be? This is a nice example of the kind of thing we need to think about when defining a measure of length for Unicode strings.
- danbruc 3y agoFour. Obviously. The more interesting question is whether the Unicode rules actually give that answer. EDIT: Just checked it using the first online tool [1] that came up and it indeed says four. So all is good. [1] https://onlinetools.com/unicode/extract-unicode-graphemes https://onlinetools.com/unicode/extract-unicode-graphemes
- masklinn 3y agoIt should be 4 as long as you count the grapheme clusters which is what e.g. Swift does (hence String#count being O(n)). In Javascript, you can get the same information through Intl.Segmenter, segments by grapheme cluster by default.
- danbruc 3y agoYou could also have it in O(1), just store and maintain it as you usually store the length in bytes or code units. If you had all your string operations like substring work with grapheme clusters by default, which might arguably make sense quite often, then that could actually be a good decision. It might even make sense to maintain a list with pointers to each grapheme cluster or of all the grapheme cluster lengths together with the actual string data. Or maybe not, would probably depend heavily on the workload.
- coding123 3y agoHonestly the "what encoding is this! UTF-8" is still the only thing we need to know. len(emoji) is still a corner case that few will care about.
- mcfedr 3y agoThat's what everyone thinks, until the user sticks an emoji in the name field
- jaza 3y agoNo emojis (or anything else remotely exotic) in names thanks. /^[A-Za-z\-' ]+$/ Users can beg and grovel at my feet for every measly character beyond that puny set that they want allowed in a name field.
- ssokolow 3y agoMy anglophone Canadian brother's name is André. Even if you're fine with alienating the ~50% of the world population using non-latin writing systems, probably best to at least stick to the stuff covered by the latin1 legacy encoding.
- tripdout 3y agoCan there be overlaps between fonts in the private use area?
- mankyd 3y agoYes. "Private" in this case means that you can't expect consistent behavior from one system to the next.
- punkbit 3y agoThe mouse cursos ir really annoying, stopped reading for that reason
- moelf 3y ago> The only modern language that gets it right is Swift: arguably not true: julia> using Unicode # for some reason HN doesn't allow emoji julia> graphemes(" ") length-1 GraphemeIterator{String} for " " help?> graphemes search: graphemes graphemes(s::AbstractString) -> GraphemeIterator Return an iterator over substrings of s that correspond to the extended graphemes in the string, as defined by Unicode UAX #29. (Roughly, these are what users would perceive as single characters, even though they may contain more than one codepoint; for example a letter combined with an accent mark is a single grapheme.)
- gwbas1c 3y agoJulia is not a major language like Swift.
- JRaspass 3y agoRaku also gets it right.
- SyrupThinker 3y agoI imagine the author would disagree with that because it does not have the “right” behavior by default. For example indexing and length of the string are done by codeunit. [1] On the other hand Rakus Str type does behave similarly to Swifts: indexing, length and iteration by grapheme; view methods for specific encodings. [2] [1]: https://docs.julialang.org/en/v1/base/strings/ https://docs.julialang.org/en/v1/base/strings/ [2]: https://docs.raku.org/type/Str#routine_chars https://docs.raku.org/type/Str#routine_chars
- WillAdams 3y agoJust had this come up at work --- needed a checkbox in Microsoft Word --- oddly the solution to entering it was to use the numeric keypad, hold down the alt key and then type out 128504 which yielded a check mark when the Arial font was selected _and_ unlike Insert Symbol and other techniques didn't change the font to Segoe UI Symbol or some other font with that symbol. Oddly, even though the Word UI indicated it was Arial, exporting to a PDF and inspecting that revealed that Segoe UI Symbol was being used. As I've noted in the past, "If typography was easy, Microsoft Word wouldn't be the foetid mess which it is."
- uxp8u61q 3y agoThat's unrelated to unicode. The checkmark symbol just isn't in the Arial font, so Word just falls back to a font that has it - Segoe UI. You've found a bug where Word still thinks it's Arial. But this is something that would happened no matter what encoding you choose for your characters.
- neerajsi 3y agoI don't know this for a fact, but it's possible that the text run is logically considered to be Arial and the fallback could be handled as just a rendering step, rather than being encoded in the document. Doing it that way could allow the text to render on different versions of Arial, some of which do have a checkbox char, at the risk of the appearance and layout changing depending on which fonts are installed.
- WillAdams 3y agoThe weird thing is, it didn't work thus for other numbers for used for this or similar characters.
- Tomte 3y ago> They will look the same (Å vs Å) No. In my browser the first A has the ring glued to it, and the second has a little gap.
- nabla9 3y ago>3 Grapheme Cluster Boundaries >It is important to recognize that what the user thinks of as a “character”—a basic unit of a writing system for a language—may not be just a single Unicode code point. Instead, that basic unit may be made up of multiple Unicode code points. To avoid ambiguity with the computer use of the term character, this is called a user-perceived character. For example, “G” + grave-accent is a user-perceived character: users think of it as a single character, yet is actually represented by two Unicode code points. These user-perceived characters are approximated by what is called a grapheme cluster, which can be determined programmatically.
- sebstefan 3y agoOh my god, is there ever anything simple about unicode
- WorldMaker 3y agoCompared to the ancient world of EBCDIC versus ASCII versus various ISO standards versus country-defined encodings versus Extended EBCDIC code pages versus Extended ASCII code pages which varied depending on operating system, nearest flag pole, network adapter, time of day, etc…: Unicode will forever be a simpler walk in the park. It's complexity is a relief compared to where we've been. It's definitely not simple, but it will forever be far simpler than what our grandmothers had to work with if they were writing international software.
- ssokolow 3y agoGive the "Indic scripts" section of https://manishearth.github.io/blog/2017/01/15/breaking-our-latin-1-assumptions/ https://manishearth.github.io/blog/2017/01/15/breaking-our-l... a read. TL;DR: Unicode is complicated because some non-Latin writing systems are complicated and those non-Latin writing systems account for over a quarter of the world's population. (They're either majority or present in India, Indonesia, Pakistan, Bangladesh, the Philippines, etc.)
- nottorp 3y ago> These user-perceived characters are approximated by what is called a grapheme cluster, which can be determined programmatically. From everything i've read or heard about unicode, "determined programmatically" is false?
- jcranmer 3y agoThere's one part of this document that I would push extremely hard against, and that's the notion that "extended grapheme clusters" are the one true, right way to think of characters in Unicode, and therefore any language that views the length in any other way is doing it wrong. The truth of the matter is that there are several different definitions of "character", depending on what you want to use it for. An extended grapheme cluster is largely defined on "this visually displays as a single unit", which isn't necessarily correct for things like "display size in a monospace font" or "thing that gets deleted when you hit backspace." Like so many other things in Unicode, the correct answer is use-case dependent. (And for this reason, String iteration should be based on codepoints--it's the fundamental level on which Unicode works, and whatever algorithm you want to use to derive the correct answer for your purpose will be based on codepoint iteration. hsivonen's article (https://hsivonen.fi/string-length/ https://hsivonen.fi/string-length/), linked in this one, does try to explain why extended grapheme clusters is the wrong primitive to use in a language.)
- pif 3y ago> An extended grapheme cluster is largely defined on "this visually displays as a single unit", which isn't necessarily correct for things like "display size in a monospace font" or "thing that gets deleted when you hit backspace." I'm sorry, but I fail to see how "This visually displays as a single unit" could ever differ from "Display size in a monospace font" or "Thing that gets deleted when you hit backspace".
- mattnewton 3y ago> Display size in a monospace font Some clusters are going to be multiple characters wide. > thing that gets deleted when you hit backspace Some clusters are meant to be composted of multiple keystrokes and a natural editing experience would allow users to delete the last stroke. Look into how Korean works.
- jcranmer 3y agoSee, e.g., https://github.com/xi-editor/xi-editor/issues/655 https://github.com/xi-editor/xi-editor/issues/655 for why backspace isn't the same as extended grapheme cluster. As for "display size in monospace font", emojis and CJK characters are usually two units wide, not one (although, to be honest, there's a fair amount of bugs in the Unicode properties that define this).
- hot_gril 3y agoThis is a lot more than the minimum that every software dev must know about Unicode. Even if you only do web frontends, you will do fine not knowing most of this. Still a nice read, though.
- pif 3y ago> The minimum every software developer must know about Unicode Just a nitpick... Once more, as it is typical on HN, web programming is confused with the entire universe of software development. There are plenty of software realms where ASCII not only is enough, but it actually MUST be enough.
- lxgr 3y agoWhat do you mean by “must be enough”? Not being able to support non-latin scripts sounds more like a limitation than a feature to me, although of course in many contexts it’s not in any individual organizations power to overcome it.
- 9dev 3y agoWell, proper Unicode support affects pretty much any area handling data about, used by, or created by, humans. That’s a pretty broad scope, and certainly wider than just web software.
- uxp8u61q 3y agoThis kind of assertiveness leads to garbage like C++ still not supporting UTF8 properly in 2023. My name contains diacritics. I am so, so, so tired of trying to work around information systems - not just web frontends - designed by people who don't care or worse, don't want to care. "Web" programmers can care all they want about Unicode, but if the backend people didn't deal properly with text encoding, then something will break no matter what. > There are plenty of software realms where ASCII not only is enough, but it actually MUST be enough. Name one.
- pif 3y ago> if the backend people didn't deal properly You are right. It's not a frontend/backend issue. It's a "for human" vs "not for human" issues. Personal names must be treated in an international-friendly manner. >> There are plenty of software realms where ASCII not only is enough, but it actually MUST be enough. > > Name one Joel himself described an example: > It would be convenient if you could put the Content-Type of the HTML file right in the HTML file itself, using some kind of special tag. Of course this drove purists crazy… how can you read the HTML file until you know what encoding it’s in?! Luckily, almost every encoding in common use does the same thing with characters between 32 and 127, so you can always get this far on the HTML page without starting to use funny letters: The content of a webpage is required to be expressed in every supported language, but the HTTP protocol must not. And it would make no sense at all to add internationalization to intra-machines protocol, where ASCII is enough and has been enough for decades. And if someone complains that ASCII only supports English, well... suck it up! I'm Italian and work in French, still I hate when a colleague sneaks in a comment not in English. Professional software development happens in English.
- throwaway_fjmr 3y agoAnd yet, many modern, recent apps can't even encode the accented European character in my given name. Sigh.
- bluecheese452 3y agoAnyone else hate titles like this? There are millions of developers working on a large variety of things. It sounds so arrogant to me.
- extraduder_ire 3y agoIt references previous famous blogposts. Much like "$THING considered harmful" titles. You'd also have a hard time working with computers in the modern day without running into unicode.
- beders 3y agoTonsky, dude. I stopped reading your article because of your little websocket experiment.
- gh0stcloud 3y agother article's background color deserves to be named: https://colornames.org/color/fddb29 https://colornames.org/color/fddb29
- kajaktum 3y agoI am torn between supporting all languages (which easily leaks into supporting emojis) versus just using the 90~ Latin characters as the lingua franca. Look, I would love to be able to read/write Sanskrit, Arabic, Chinese, Japanese etc and share those content and have everyone render and see the same thing. The problem is that I feel like most of these are: 1. a kind of an open problem 2. very subjective 3. very, very subjective as what you is mostly dictated by the implementation (fonts) For example, why does a gun emoji looks like water gun? Why is the skull-and-crossbones symbol looks so benign. In fact, it is often used as a meme (see deadass :skull:) Why is the basmala a single "character"? In my opinion, people should just learn how to use kaomoji. Granted, kaomojis rely on a lot more than the Latin characters but it is at least artful, skillfull and a natural extension of the "actual" languages. > inb4 languages evolves Yes, but it mostly happens naturally. I feel like what happens today mostly happens at the whim of a few passionate people in the standard.
- zzo38computer 3y ago> I am torn between supporting all languages (which easily leaks into supporting emojis) versus just using the 90~ Latin characters as the lingua franca. I don't want to support emoji either (and, I don't want emoji on my computer), although in some cases, if it is really necessary to be supported, they could be implemented just as text characters instead of as colourful emoji, anyways. For many purposes (e.g. computer codes) ASCII is good enough (and actually even can be better since it avoids the security problems of using Unicode). (Sometimes, character sets other than ASCII can be used, e.g. APL character set for APL programming.) > Look, I would love to be able to read/write Sanskrit, Arabic, Chinese, Japanese etc I also would, but Unicode is bad enough that I would use other ways of doing such a thing when possible (even writing my own programs, etc). (If a program insists on Unicode, I might just use ASCII only anyways, or write my own program) Not everyone necessarily need to see the same thing (if it is a text, rather than pictures of the text), although, the suitable character sets for that language which can be in use (and with fix pitch if necessary, etc), to auto select a suitable fonts for your computer by the reader's preference. So, I prefer to support all languages (where applicable; sometimes it isn't), without using Unicode.
- Nevermark 3y agoWhat an interesting mess! It occurs to me that a canonical semantic representation of all known (extracted) language concepts would be useful too. Now that we have multi-language LLM's it would be an interesting challenge to create/design a canonical representation for a minimum number of base concepts, their relations and orthogonal "voice" modifiers, extracted from the latent representations of an LLM across a whole training set, over all training languages. While the best LLMs still have complex reasoning issues, their understanding of concepts and voice at the sentence level is highly intuitive and accurate. So the design process could be automated. The result would be a human language agnostic, cross-culture concept inclusive, regularized & normalized (relatively speaking) semantic language. Call it SEMANTICODE. We need to get this right, using one standard LLM lineage, before the Unicode people create a super standard that spans 150 different LLM's and 150 different latent spaces! :O Stability between updates would be guaranteed by including SEMANTICODE as a non-human language in training of future LLM's. Perhaps including a (highly) pre-normalized semantic artificial language would dramatically speed up and reduce the parameter count needed for future multi-language training?* Then LLMs could use SEMANTICODE talk to each other more reliably, efficiently, and with greater concept specificity than any of our single languages.
- chx 3y ago"roll your own" Rather not. It takes an incredible amount of work to get it right. Just stick to ICU.
- w10-1 3y agoA real question is why IBM, Apple, and Microsoft poured millions into developing the unicode standard instead of treating character encoding like file formats as a venue for competition. IBM and Apple in the early 1990's combined in Taligent to try to beat MS NT, but failed. But a lot of internationalization came out of that and was made open, at the perfect time for Java to adopt it. Interestingly it wasn't just CJK but Thai language variants that drove much of the flexibility in early unicode, largely because some early developers took a fancy to it. When you look at the actual variety in written languages, Unicode grapheme/code-point/byte seems rather elegant. We're in the early days of term vectors, small floats, and differentiable numerics (not to mention big integers). Are lessons from the history of unicode relevant?
- preciousoo 3y agoYou can ask why they didn’t do the same for networking and serial protocols too.
- deleted 3y ago[deleted]
- thyselius 3y agoWonderful to learn more about Unicode. Does anyone know how to write a function (preferably in swift) to remove emoji? This is surprisingly hard (if the string can be any language, like English or Chinese). There’s been multiple attempts on Stackoverflow but they’re all missing some of them, as Unicode is so complex.
- fiedzia 3y agoI haven't tried but use libicu (icu). Split text into graphemes and remove anything starting with codepoints that has Zsey script. There should be swift bindings.
- favorited 3y agoHere's a 1-liner, producing the string "text 0123 漢字": `String("text EMOJI 0123 漢字".unicodeScalars.filter({ !$0.properties.isEmojiPresentation }))` (I've had to substitute EMOJI for a smiley face, because HN is bad at text encoding.)
- nottorp 3y ago> Since everybody in the world agrees on which numbers correspond to which characters, and we all agree to use Unicode, we can read each other’s texts. Hmm? I thought some code points combine to create a character. Even accented latin ones can be like that. Also we need to agree on what is a character.
- JohnFen 3y ago> Also we need to agree on what is a character. Indeed. I used to think I knew what a character was until Unicode came around. Now I genuinely don't know with any real certainty.
- aembleton 3y agoIn Java/Kotlin, I've found this Grapheme Splitter library to be useful: https://github.com/hiking93/grapheme-splitter-lite https://github.com/hiking93/grapheme-splitter-lite
- kipcole9 3y ago> The only modern language that gets it right is Swift: Elixir too: Interactive Elixir (1.15.4) - press Ctrl+C to exit (type h() ENTER for help) iex(1)> String.length "ẇ͓̞͒͟͡ǫ̠̠̉̏͠͡ͅr̬̺͚̍͛̔͒͢d̠͎̗̳͇͆̋̊͂͐" 4
- rkagerer 3y ago"the definition of graphemes changes from version to version" In what twisted reality did someone think this a good idea? Doesn't it go against the whole premise of everyone in the world agreeing on how to represent a meaningful unit of text? "What’s sad for us is that the rules defining grapheme clusters change every year as well. What is considered a sequence of two or three separate code points today might become a grapheme cluster tomorrow! There’s no way to know! Or prepare!" "Even worse, different versions of your own app might be running on different Unicode standards and report different string lengths!"
- rkagerer 3y agoI can sympathize why some programmers would prefer to stick their heads in the sand and stick to ASCII.
- Dwedit 3y agoReload with Javascript disabled to remove the distracting fake mouse pointers.
- diego_sandoval 3y agoExtended Grapheme Cluster should be understood as Extended (Grapheme Cluster) or as (Extended Grapheme) Cluster?
- ssokolow 3y ago"Extended (Grapheme Cluster)". The .graphemes() method in Rust's unicode-segmentation crate takes an is_extended boolean as an argument and, if you set it to false, you're iterating legacy grapheme clusters.
- wickedsickeune 3y agoI'm sorry but the website design is extremely distracting. The mouse pointers at least are easy to delete with the inspector; The background color is not the best choice for reading material, but the inexcusable part is the width of the content. This content must be really awesome for someone to go through the trouble of interacting with such a site.
- deleted 3y ago[deleted]
- overflyer 3y ago[flagged]
- hyggetrold 3y agoIs there a way to read this with the mouse cursors disabled? It seems like great content but all the movement on the page is way too distracting. EDIT: I've never been downvoted for asking a question before. Weird, but okay.
- rdtsc 3y ago> The only modern language that gets it right is Swift: print("...".count) // => 1 And Erlang/Elixir! I guess they are not "cool" enough. But they correctly interpret that as one grapheme cluster. % erl +pc unicode > string:length("..."). 1 (... here is the U+1F926 U+1F3FB U+200D U+2642 U+FE0F emoji)
- wodow 3y agoThe author does refer to Elixir further down: > UPD: Erlang/Elixir seem to be doing the right thing, too.
- rdtsc 3y agoUpdated a day after it was mentioned here :-) https://github.com/tonsky/tonsky.me/commit/5fbbb373025be375841cbb769dced514dec37312 https://github.com/tonsky/tonsky.me/commit/5fbbb373025be3758...
- zackmorris 3y agoThe only modern language that gets it right is Swift: Apple did a fairly good job with unicode string handling starting in Cocoa and Objective-C, by providing methods to get the number of code points and/or bytes: https://stackoverflow.com/questions/15582267/cfstring-count-of-characters-not-code-points-in-a-string/15582268#15582268 https://stackoverflow.com/questions/15582267/cfstring-count-... I feel that this support of both character count and buffer size in bytes is probably the way to go. But Python 3 went wrong by trying to abstract it away with encodings that have unintuitive pitfalls that broke compatibility with Python 2: https://blog.feabhas.com/2019/02/python-3-unicode-and-byte-strings/ https://blog.feabhas.com/2019/02/python-3-unicode-and-byte-s... There's also the normalization issue. Apple goofed (IMHO) when they used NFD in HFS+ filenames while everyone else went with NFC, but fixed that in APFS: https://unicode.org/faq/normalization.html https://unicode.org/faq/normalization.html https://medium.com/@sthadewald/the-utf-8-hell-of-mac-osx-feef5ea42407 https://medium.com/@sthadewald/the-utf-8-hell-of-mac-osx-fee...
- WalterBright 3y agoQuotes from the article illustrating what a train wreck Unicode has become: "The problem is, in Unicode, some graphemes are encoded with multiple code points!" "An Extended Grapheme Cluster is a sequence of one or more Unicode code points that must be treatead as a single, unbreakable character." "Starting roughly in 2014, Unicode has been releasing a major revision of their standard every year." "Å" === "Å" "Å" === "Å" "Å" === "Å" What do you get? False? You should get false, and it’s not a mistake. "That’s why we need normalization." "Unicode is locale-dependent" The article forgot one: characters that switch presentation to right-to-left.
- penguin_booze 3y agoI knew that domain, so I had sunglasses at hand before opening the page!
- user3939382 3y agoI once bought an O'Reilly book on encoding. It was like 2000 pages. I never read it, that was about 15 years ago. My take away is that encoding is really complex and I just kind of pray it works which most of the time it does.
- phforms 3y agoRegarding UTF-8 encoding: “And a couple of important consequences: - You CAN’T determine the length of the string by counting bytes. - You CAN’T randomly jump into the middle of the string and start reading. - You CAN’T get a substring by cutting at arbitrary byte offsets. You might cut off part of the character.” One of the things I had to get used to when learning the programming language Janet is that strings are just plain byte sequences, unaware of any encoding. So when I call `length` on a string of one character that is represented by 2 bytes in UTF-8 (e.g. `ä`), the function returns 2 instead of 1. Similar issues occur when trying to take a substring, as mentioned by the author. As much as I love the approach Janet took here (it feels clean and simple and works well with their built-in PEGs), it is a bit annoying to work with outside of the ASCII range. Fortunately, there are libraries that can deal with this issue (e.g. https://github.com/andrewchambers/janet-utf8 https://github.com/andrewchambers/janet-utf8), but I wish they would support conversion to/from UTF-8 out of the box, since I generally like Janet very much. One interesting thing I learned from the article is that the first byte can always be determined from its prefix. I always wondered how you would recognize/separate a unicode character in a Janet string since it may have 1-4 bytes length, but I guess this is the answer.
- pornel 3y agoYou CAN'T do any of these things in Unicode in general, in all of its encodings. There's no random access in Unicode. It's a stateful system that requires linear scan.
- russellbeattie 3y agoWe need, desperately and without question, two Unicode symbols for bold and italic. These are part of language and should not be an optional proprietary add on that can be skipped or deleted from text. We've been using the two "formats" to convey important information since the sixteenth century!!! It boggles my mind that we can give flesh tone to emojis, yet not mark a word as bold or italic. It makes zero sense. Especially how easy it would be to implement. It would work exactly the same way: Letters following the mark would be formatted as bold or italic until a space character or equivalent.
- cryptonector 3y ago> Another unfortunate example of locale dependence is the Unicode handling of dotless i in the Turkish language. This isn't quite Unicode's fault, as the alternative would be to have two codepoints each for `i` and `I`, one pair for the Latin versions and one for the Turkish versions, and that would be very annoying too. Whereas the Russian/Bulgarian situation is different. There used to be language tags in Unicode for that, but IIRC they got deprecated, and maybe they'll have to get undeprecated.
- TacticalCoder 3y ago> For example, é (a single grapheme) is encoded in Unicode as e (U+0065 Latin Small Letter E) + ´ (U+0301 Combining Acute Accent). Two code points! It's a poor and misleading example for it is definitely not how 'é' is encoded in 99.999% of all the text written in, say, french out there (french is the language where 'é' is the most common). 'é' is U+00F9, one codepoint, definitely not two. Now you could say: but it is also the two codepoints one. But that's precisely what makes Unicode the complete, total and utter clusterfuck that it is. And hence even an article explaining what every programmer should know about Unicode cannot even get the most basic example right. Which is honestly quite ironic.
- wgjordan 3y agoThe author explains normalization in its own entire section several paragraphs later (Why is "Å" !== "Å" !== "Å"?).
- spacechild1 3y agoNext time read the whole article before accusing the author of incompetence! However, the author could have added a small note, e.g. "(Unicode normalization will be convered in a later section.)", to prevent knowledgable readers from rage quitting :)
- ninkendo 3y ago> Unicode the complete, total and utter clusterfuck that it is. Yikes, does it really deserve that much derision? They’re trying to standardize all written human language here. I think they’ve done a fantastic job. Pre-Unicode you had to worry about what code page a document had, and computers from different countries couldn’t interoperate. The work the consortium does is hugely important, and every decision has extremely complex tradeoffs. Composed characters makes a lot of sense, and there’s a lot of case to be made that it was the right call. The attitude of “this one thing I don’t like makes the whole thing a complete clusterfuck” is something I wish fewer engineers would have.
- rstuart4133 3y ago> Yikes, does it really deserve that much derision? To my simple mind it had one job: allocate every grapheme a number (code point). Had it done that, the 1/2 of the article warning you about the difficulty of iterating and modifying code points would have disappeared. But I guess it had a 2nd job: create a way of representing those numbers. The obvious way, u32, was difficult for ASCII users swallow as it quadrupled the space used for a string. The solution we settled on, UTF-8 didn't come form Unicode (or ISO). It came from Ken Thompson (the Ken Thompson, who created B, the predecessor of C), when he tired to make something workable for C. Unicode was the entity that ballsed up both of those tasks. It was a fork of ISO 10646. It's main contribution over 10646 was UCS-2 - ie 16 bits per character. That decision was so bad it had to be abandoned. Later they introduced the grapheme clusters rather than allocating a separate code point for each variant. I have no idea why, as it makes the programmers task far harder. Maybe they ran out of code points. How could they possibly run out of code points, given U32 has 4 billion of them and UTF-8 could potentially have more? Because they had to kludge their way around the USC-2 mistake to create UTF-16, and it's limited 1 million. Which leads us to the one thing in the article I disagree with: > The only downside of UTF-16 is that everything else is UTF-8, so it requires conversion every time a string is read from the network or from disk. No, that's not the only downside. There is one more: USC-2 / UTF-16 has the endianness problem. A 16 bit value needs two bytes to represent it, and you can write two bytes to storage in two ways - little endian or big endian. They didn't specify, so the same string can have two different representations on disk. They added the infamous BOM markers to distinguish between them. I could go on, but colour me singularly unimpressed with this mob.
- titzer 3y ago> The problem is, you don’t want to operate on code points. A code point is not a unit of writing; one code point is not always a single character. What you should be iterating on is called “extended grapheme clusters”, or graphemes for short. It's best to avoid making overly-general claims like this. There are plenty of situations that warrant operating on code points, and it's likely that software trying and failing to make sense of grapheme clusters will result it in a worse screwup. Codepoints are probably the best default. For example, it probably makes the most sense for programming languages to define strings as arrays of code points, and not characters or 16-bit chunks or an encoding, or whatever.
- Dylan16807 3y agoSituations such as? Sometimes editing wants to go inside clusters but that's not code-point based either. I'd say that in a big majority of situations, code that is indexing an array with code points is either treating the indexes as opaque pointers or is doing something wrong.
- hgs3 3y ago> There are plenty of situations that warrant operating on code points Absolutely correct. All algorithms defined by the Unicode Standard and its technical reports operate on the code point. All 90+ character properties defined by the standard are queried for with the code point. The article omits this information and ironically links to the grapheme cluster break rules which operate on code points.
- Dylan16807 3y agoThe article doesn't say not to use code points, it says you should not be iterating on them. Very rarely will you be implementing those algorithms. And if you're looking at character properties, the article says you should be looking at multiple together, which is correct.
- hgs3 3y ago> And if you're looking at character properties, the article says you should be looking at multiple together, which is correct. I don't see where the article mentions Unicode character properties [1]. These properties are assigned to individual characters, not groups of characters or grapheme clusters. > Very rarely will you be implementing those algorithms. True, but character properties are frequently used, i.e. every time you parse text and call a character classification function like "isDigit" or "isControl" provided by your standard library you are in fact querying a Unicode character property. [1] https://unicode.org/reports/tr44/#Properties https://unicode.org/reports/tr44/#Properties
- charcircuit 3y agoThe number of graphene clusters in a string depend on the font being used. The length of a string should be the number of code points because that is not length specific. Better yet, there shouldn't be a function called length.
- loeg 3y agoPretty clearly, "every software developer" doesn't need to understand Unicode with this level of familiarity, much like "every programmer" doesn't need to know the full contents of the 114 page Drepper paper. For example, I work on a GUID-addressed object store. Everything is in term of bytes and 128-bit UUIDs. Unicode is irrelevant to everyone on my team, and most adjacent teams. There is lots of software like this.
- koreth1 3y agoGlad I'm not the only one who was irked by this, and I do need to know a lot about Unicode for my job! I believe there actually are topics that every software developer ought to know something about, but this isn't one of them. My list would be things more like, the difference between a constant-time algorithm and a quadratic-time one.
- PeterisP 3y agoThere seems to be quite large segments of developers working on functionality which "handles text" as immutable whole blobs, in which case one really doesn't have to know anything about unicode. However, as soon as you want to look into the text contents of the objects of that object store and handle parts of it, even in the very simplest way (e.g. checking whether the stored object contains some character, or whether two text messages stored in that object store are the same) then you can't treat them as bytes anymore, and all the concerns listed in this article suddenly become relevant for your team.
- loeg 3y agoYes, exactly.
- fomine3 3y agoEvery programmer don't have to remember all things in this article, but they should remember that Unicode (or text system in the wild) is actually complex so they should research as needed.
- 3y ago
- kazinator 3y agoIf you have to recognize a grapheme cluster, it will be easier to do that from a sequence of code points, than from UTF-8. It's like saying that we don't need to tokenize, because you never want to deal with tokens anyway, but phrase structures! Mmkay, whatever ...
- thefringthing 3y ago> Unicode is a standard that aims to unify all human languages, both past and present, and make them work with computers. This is doubly wrong. First, it conflates languages and writing systems. Malay and English use the same writing system but are different languages. American Sign Language is a language, but it has no standard or widely-adopted writing system. Hakka is a language, but Hakka speakers normally write in Modern Standard Mandarin, a different language. Second, it's not that case that Unicode aims to encode all writing systems. For example, there are many hobbyist neographies (constructed writing systems) which will not be included in Unicode.
- bmicraft 3y agoIf you consider that it just "aims to" and makes no claim of succeeding to unify all languages, it isn't that wrong
- rcaught 3y ago> Second, it's not that case that Unicode aims to encode all writing systems. For example, there are many hobbyist neographies (constructed writing systems) which will not be included in Unicode. Doesn't the "private use space" technically satisfy this?
- kalinf 3y ago> Among them is assigning the same code point to glyphs that are supposed to look differently, like Cyrillic Lowercase K and Bulgarian Lowercase K (both are U+043A). This is nonsense, Bulgaria has been using the Cyrillic alphabet sinse its creation in … Bulgaria! What you’ve shown is two different fonts, and both renderings are perfectly fine in Bulgaria. Read up more about it on wikipedia: https://en.wikipedia.org/wiki/Bulgarian_alphabet https://en.wikipedia.org/wiki/Bulgarian_alphabet
- deleted 3y ago[deleted]
- gorgoiler 3y agoWith the benefit of hindsight, would we include the error detection bits of UTF8 if we could choose not to?
- ssokolow 3y agoYes. Give https://www.youtube.com/watch?v=_mZBa3sqTrI https://www.youtube.com/watch?v=_mZBa3sqTrI a watch... especially the "Oh my God! We've been hacked!" part at 36:20. TL;DR: They had a transient glitch in their network switch and, because Windows uses UTF-16 when sending remote event logs over the wire, whenever it dropped a single byte, it had the effect of swapping the endianness of the messages, resulting in scary Chinese text in the logs. You could get the same effect by naively applying byte-wise processing to UTF-16 or UTF-32, or having an off-by-one error. UTF-8 is self-synchronizing so one-byte errors like that only lose you one character, rather than corrupting the entire stream going forward.
- everyone 3y agoI'm always gonna point out these overly broad titles assuming "every software developer" is some kind of internetty web dev type. I'm a game dev, I try and never touch strings at all, they are a nightmare data type. Strings in a game are like graphics or audio assets, your game might read them and show them to player, but they should never come anywhere near your code or even be manipulated by it. I dont need to know any of that stuff about Unicode.
- Terr_ 3y ago> normalization A quick war-story on this: We had a system which was taking web-user input for human names, and then some of it had to be sanitized for a crappy third-party system. However some of the names were getting mangled in unexpected ways. One of the (multiple) issues was that we were sometimes entirely dropping accented characters even when a good alternative existed. This occurred when we were getting "é" (U+00E9) instead of "é" (U+0065 U+0301), a regular letter E plus a special accent modifier. By forcing the second form (D normalization) we were able to strip "just the accents" and avoid excessively-wrong names. Going further with K+D normalization, weird stuff like "⑧" (letter 8 in a circle) becomes a regular number 8.
- ssokolow 3y agoGive https://manishearth.github.io/blog/2017/01/15/breaking-our-latin-1-assumptions/ https://manishearth.github.io/blog/2017/01/15/breaking-our-l... a read and, ideally, the other things it mentions like https://eev.ee/blog/2015/09/12/dark-corners-of-unicode/ https://eev.ee/blog/2015/09/12/dark-corners-of-unicode/ and https://manishearth.github.io/blog/2017/01/14/stop-ascribing-meaning-to-unicode-code-points/ https://manishearth.github.io/blog/2017/01/14/stop-ascribing.... (Among other things, it points out that doing that to non-Latin text is liable to change pronunciations and meanings in other languages. For example, some languages use diacritics for voiced/unvoiced indication where your "normalization" could do things like "tick→dick" or "did→tit".) (Did you ever notice that? B/P, D/T, V/F, G/K, J/CH, and Z/S form voiced/unvoiced pairs that could have been indicated with a single letter and a diacritic. Same mouth behaviour. It's just a question of whether you engage your vocal cords.)
- throwaway2037 3y agoIt sounds like a generic length function in Unicode in 2023 is no longer a good idea. These articles complaining about the variety of lengths in Unicode are annoying at this point. Pretty much all of them can be summed up as, "Well, it depends." And, that isn't wrong. But nerds love to argue until they are blue in the face about the One Correct Answer. Sheesh. This is the most interesting comparison article I have seen in years about Unicode processing in C++: https://thephd.dev/the-c-c++-rust-string-text-encoding-api-landscape https://thephd.dev/the-c-c++-rust-string-text-encoding-api-l... The author is also the lead on an open source C++ Unicode library called ztd.txt: https://github.com/soasis/text https://github.com/soasis/text
- redder23 3y agoIf just for the fact that it annoys people I love the mouse cursor idea. But I also find it technically interesting. Is some kind of consent legally needed per GDPR or something for this? I for sure is tracking, literally. And a website has to ask to set cookies ...
- ssokolow 3y agoCookie consent is only necessary if you're sharing it with others (eg. ad networks, Google Analytics, etc.) or using it for "non-essential" functions (again, stuff like analytics). Sites just don't want the general public to realize that. As for the mouse cursors, I don't think they qualify as personal information under the GDPR, but IANAL.
- extraduder_ire 3y agoIt's needed when you collect the data. Cursor position is being used directly to make the service "function" though. Even if the function it enables is pretty novel. This is entirely debatable, and probably matters less than showing were your cursor is in a google doc. Regardless, I don't think it matters since the author is not in the EU.
- tootie 3y agoOne of my favorite interview questions is simply "What's the difference between Unicode and utf-8?" I feel like that's pretty mandatory knowledge for any specialty but it doesn't get answered correctly very often.
- deleted 3y ago[deleted]
- ssokolow 3y ago> Before comparing strings or searching for a substring, normalize! ...and learn about the TR39 Skeleton Algorithm for Unicode Confusables. Far too few people writing spam-handling code know about that thing. (Basically, it generates matching keys from arbitrary strings so that visually similar characters compare identical, so those Disqus/Facebook/etc. spam messages promoting things like BITCO1N pump-and-dumps or using esoteric Unicode characters to advertise work-from-home scams will be wasting their time trying to disguise their words.) ...and since it's based on a tabular plaintext definition file, you can write a simple parser and algorithm to work it in reverse and generate sample spam exploiting that approach if you want. https://www.unicode.org/Public/security/latest/confusables.txt https://www.unicode.org/Public/security/latest/confusables.t... > and CD-ROM! I think you mean Microsoft Windows's Joliet extensions to ISO9660 which, by the way, use UCS-2, not UTF-16. (Try generating an ISO on Linux (eg. using K3b) with the Joliet option enabled and watch as filenames with emoji outside the Basic Multilingual Plane cause the process to fail.) The base ISO9660 filesystem uses bytewise-encoded filenames.
- dystroy 3y agoBut not all normalizations are done to fight spam, not all of them should be interested in visual similarity. I normalize strings in searches not because of bad intents but because for all user related purposes "Comunicações" and "Comunicações" are the same, their different encodings being more of an accident.
- ssokolow 3y ago*nod* ...and stemming is that taken to a greater extreme. I was just pointing out that Unicode itself has various forms of normalization and normalization-adjacent functionality that people are far too unaware of.
- Avlin67 3y agoFirefox 120.0a1 has difficult time displaying this page
- shellmachine 3y ago$ ping6 tonsky.me ping6: no address associated with name $ What every software developer must know about IPv6 in 2023 (still no excuses!).
- denton-scratch 3y ago> Hell, you can’t even spell café, piñata, or naïve without Unicode. I must have missed something. All of those symbols are present in Extended ASCII (i.e. 8-bit).
- extraduder_ire 3y agoExtended ascii is a bodge, and requires you to set a code page to pick the right set of "extended" characters. Unicode is also a superset of ascii though, so that sentence is right, on a technicality.
- racl101 3y agoThe fuck is up with the cursor nonsense? I would've read this thing if it wasn't for that.
- velox_neb 3y agoErrata: In the table under "How many bytes are in UTF-8?", bottom row, "10000" should be "100000".
- garfieldnate 3y agoWhat's the point of having a separate codepoint for the Angstrom if it's specified to normalize back to the regular "capital A with ring above" codepoint anyway?
- tqwhite 3y agoBest. Explanation. Ever.
- honkyponky 3y agohttps://worldswritingsystems.org/ https://worldswritingsystems.org/
- deleted 3y ago[deleted]
- ycomfakeuser 3y agoThere is another 'modern' language that does utf8 right and has done it right for a long time. I know it's mostly fallen out of favour, but we're still out here: Perl. $ perl -wle 'use utf8; print length("");' 1 Without use utf8: $ perl -wle 'print length("");' 3 It's funny: after Perl fell out of favour, is when it got all its best stuff. It's still my preferred language for just about everything.
- ycomfakeuser 3y ago[flagged]
- khakiem 3y ago> The only modern language that gets it right is Swift The only modern language ((he knows of)) that gets it right to be precise. Ruby also gets it right. ``` [1] pry(main)> "".size => 1 ```
- elcaro 3y agoAnother Unicode article that mentions Swift, but not Raku :( Raku's Str type has a `.chars` method that counts graphemes. It has a separate `.codes` method to count codepoints. It also can do O(1) string indexing at the grapheme level. That Zalgo "word" example is counted as 4 chars, and the different comparisons of "Å" are all True in Raku. You can argue about the merits of it's approach (indeed several commenters here disagree that graphemes are the "one true way" to count characters), but it feels lacking to not at least _mention_ Raku when talking about how different programming languages handle Unicode.
- makeworld 3y agoReally great article. Hitting all the points I would expect.
- dekken_ 3y agoAm I supposed to hate this website, cause I kinda do
- anymouse123456 3y agoFWIW - Right Click, Inspect. There's a div with an attribute, "pointers" in the body root. Deleting that makes the while thing a lot less stressful.
- MrResearcher 3y agouBlock Origin -> Disable Javascript Problem solved!
- bqmjjx0kac 3y agoThe mustard background with black text is harsh on the eyes.
- hoseja 3y agohttps://tonsky.me/blog/unicode/overview@2x.png https://tonsky.me/blog/unicode/overview@2x.png Wow, what abominable mix of decimal and hexadecimal.
- Karellen 3y agoWhere are the decimal numbers in that image?
- morelisp 3y agoWhat comes after 9FFFF?
- Karellen 3y agoGood catch doh.
- bajsejohannes 3y agoIt goes 90000..9FFFF then 100000..10FFFF. The latter should have been A0000..AFFFF. So the author is using hex for the last four digits and decimal for the remaining ones.
- tonsky 3y agooops :) fixed, thanks!
- davidham 3y agoIs it just me, or is anyone else seeing what looks like the mouse pointer of everyone else reading the page, like 1,000 little ants on the screen
- lifeinthevoid 3y agoyup, pretty annoying
- toastercat 3y agoAnytime tonsky's site gets posted here, I'm reminded by how awful it is, which is ironic given his UI/UX background. The site's lightmode is a blinding saturated yellow, and if you switch into darkmode, it's an even less readable "cute" flashlight js trick. I don't know why he thought this was a good idea. Thank god for Firefox reader mode.
- ericmcer 3y agoI don't think he added moving cursors all over the page because he thought it was good UI/UX, he knows what he is doing.
- permo-w 3y ago>That gives us a space of about 11 million code points. About 170,000, or 15%, are currently defined. An additional 11% are reserved for private use. The rest, about 800,000 code points, are not allocated at the moment. They could become characters in the future. 1.1 million?
- run414 3y agoYeah, the author's numbers are off by a "0". It should be "1,700,000" and "8,000,000".
- qwerty456127 3y ago> The rest, about 800,000 code points, are not allocated at the moment. They could become characters in the future. Why is Tengwar still not in Uniclde officially? What's the problem with it?
- WorldMaker 3y agoThe problem with Tengwar (and Klingon) is the problem with a lot of pop culture right now: copyright. The Tolkien Estate still exists and still litigiously upholds what it can of their copyright terms. CBS Viacom (Paramount) still claim a copyright interest in all the written forms of Klingon. Copyright is not technically violated simply by encoding the characters into a plane such as one of Unicode's, that's an easy open and shut fair use, but Unicode principals have stated they don't want to pass on the copyright burden to font authors either, which would be sued if they tried to paint some of those characters. (Why encode something that fonts aren't allowed to produce?) That should also be fair use, but the law is complicated and copyright still so often today leans in favor of the Estates and major Corporations rather than fair use and the public commons. (ETA: I'm hugely in favor that "conlang", constructed language, scripts such as these should be encoded by Unicode. I wish someday we fix the copyright problems of them.)
- teddyh 3y agoTengwar is in the Under-ConScript Unicode Registry: <https://www.kreativekorp.com/ucsur/ https://www.kreativekorp.com/ucsur/>
- qwerty456127 3y agoThe ConScript Unicode Registry is a volunteer project to coordinate the assignment of code points in the Unicode Private Use Areas (PUA). Why does tengwar have to be in the PUA, why not make it a first-class charset? It's not just a minor conlang a small group of geeks invented on a weekend, it's a well-established piece of the modern culture, isn't it?
- ssokolow 3y ago
- francisofascii 3y agoI enjoyed how the timeline graphic included Joel's article. Because my first thought was hey, isn't this the same title.
- rurban 3y agoThe Why is "Å" !== "Å" !== "Å"? section still strikes me as wrong. The strings are equal even when the representations differ.
- nextaccountic 3y agoThey are logically equal (that is, they represent the same text in an abstract way), but computing this equality in practice is expensive, because you first need to normalize the strings then compare. Most languages, when comparing strings, skip the normalization and just compare string bytes as is (or, if the string is interned, compare just the pointer)
- rurban 3y agoYou can easily do the comparison dynamically with checking for combining marks, and then do the proper lookup. No need to normalize everything, or even store the normalized variant. Though in a filesystem or username lookup you would only store it normalized.
- bajsejohannes 3y agoI just not sure why they put in the "Angstrom symbol" to begin with. If you do, then why isn't the "meter symbol" (m) also represented? Fortunately, it seems like it's marked as deprecated: https://en.wikipedia.org/wiki/Angstrom#Symbol https://en.wikipedia.org/wiki/Angstrom#Symbol
- jcranmer 3y ago> I just not sure why they put in the "Angstrom symbol" to begin with. Frequently, the answer to this is "some obscure character set had this as a distinct symbol." In this case, blame the Japanese: https://en.wikipedia.org/wiki/JIS_X_0208 https://en.wikipedia.org/wiki/JIS_X_0208 Which is why there's an 'mm' and 'cm' and other random symbols: https://www.compart.com/en/unicode/block/U+3300 https://www.compart.com/en/unicode/block/U+3300
- qwerty456127 3y ago> People are not limited to a single locale. For example, I can read and write English (USA), English (UK), German, and Russian. Which locale should I set my computer to? Ideally - the "English-World" locale is supposedly meant for us, cosmopolitans. It's included with Windows 10 and 11. Practically, as "English-World" was not available in the past (and still wasn't available on platforms other than Windows the last time I checked), I have always been setting the locale to En-US even though I have never been to America. This leads to a number of annoyances though. E.g. LibreOffice always creates new documents for the Letter paper format and I have to switch it to A4 manually every time. It's even worse on Linux where locales appear to be less easy to customize than in Windows. Windows always offered a handy configuration dialog to granularly tweak your locale choosing what measures system you prefer, whether your weeks begin on sundays or mondays and even define your preferred date-time format templates fully manually. A less-spoken about problem is Windows' system-wide setting for the default legacy codepage. I happen to use single-language legacy (non-Unicode) apps made by people from a number of very different countries. Some apps (e.g. I can remeber the Intel UHD Windows driver config app) even use this setting (ignoring the system locale and system UI language) to detect your language and render their whole UI in it. > English (USA), English (UK) This deserves a separate discussion. I doubt many English speakers (let alone those who don't live in a particular anglophone country) care to distinguish between English dialects. To us presence of a huge number of these (don't forget en-AU, en-TT, en-ZW etc - there are more!) in the options lists brings only annoyance, especially when one chooses some non-US one and this opens another can of worms. By the way I wonder how do string capitalization and comparision functions manage to work on computers of people who use both English and Turkish actively (Turkish locale distinguishes between dotted and undotted İ).
- DoughnutHole 3y ago> I doubt many English speakers care to distinguish between English dialects It's worthwhile purely for the sake of autocorrect/typo highlighting in text-editing software. I don't miss the days of spelling a word correctly in my version of English but still being stuck with the visual noise of red highlighting up and down the document because it doesn't conform to US English.
- Dudester230602 3y agoGuys, you don't need to know that crap.
- nwellnhof 3y agoUnicode is a total mess. In a sane system, "extended grapheme clusters" would equal "codepoints" and it wouldn't make a difference for 99% of languages. Now we ended up with grapheme clusters, normalization, decomposition, composition, Zalgo text, etc. But instead of deprecating this nonsense, Unicode doubled down with composed Emojis.
- eviks 3y agoBut precomposing all the potential combinations is less sane than the current mess (and you can outlaw Zalgo in the standard if you think it's a serious issue) Also, the % should measure people, not languages, that would greatly decrease the imaginary 99%
- jetbalsa 3y agoI feel its the same as with any long standing computer system we have today. It was designed as more and more of the world came online and all the growing pains it came with. Could it be built from scratch today better? Yes. Will it? No. I suspect it will be around long after we are all dead. Same with IPv4 :V
- magicalhippo 3y agoFor most software it doesn't really matter either. I've written unicode-aware software for over a decade, doing a wide variety of programs, and I've never had to bother with all that mess. If I'm parsing strings I'm looking for stuff in the 7-bit ASCII range which maps neatly onto the Unicode representations, and so I just need to take care to preserve the rest. The only trouble I've had is that a lot of programmers haven't learned, or don't get, that text encoding is a thing and that it needs to be handled. So they'll hand me an XML they claim is UTF-8 encoded, except that XML header was just copypasta and the actual XML document is encoded in some other system encoding like Windows-1252. Or worse, a mix of both.
- hot_gril 3y agoHonestly I like ipv4 better than v6. I like having a NAT and easy addresses like 192.168.1.3 instead of fe80::210:5aff:feaa:20a2. They didn't need to mess with those things just to expand the address space, like how utf8 didn't require remapping ASCII.
- alexmolas 3y agoI tried to read the articles since it seemed interesting. After exactly 30 seconds trying it I had to leave the page. Impossible to read more than two sentences with all those pointer moving there - and for a folk with ADHD even more difficult. Sorry, but I couldn't make it :(
- anthk 3y agoUse the reader mode. Or if you are under GNU/Linux, use Links/Lynx.
- AlphaCerium 3y agoNot everyone runs Linux, and not every browser has a reader mode. This should not be the solution. There should definitely be an option to disable all these features, especially the dark mode toggle, that one's a fun premise, but horrific for usability.
- anthk 3y agoTrue; but Links/Lynx exists for Windows, too. Or Netsurf. At least there are alternatives to choose. But you are right, the web sucks.
- Maken 3y agoFortunately you didn't try the dark theme.
- ilyt 3y ago> Unicode is locale-dependent Well, there is a new fact that I learned and immediately hated. The fuck were authors thinking... I am now firmly convinced people developing unicode hate developers. I suspected it before just due to how messy it was (same character having different encodings ? Really ? Fuck you), but this cements it.
- JohnFen 3y ago> people developing unicode hate developers Or at least they have a vicious indifference to us. Unicode is a nightmare.
- wffurr 3y agoYeah this is a big problem for me right now trying to pick fonts and characters for CJK. I have a bunch of bugs to fix that will require sending the locale down to the text itemization code.
- zajio1am 3y agoUnicode is not locale-dependent, just mapping from graphemes to (font) glyphs is locale/font dependent.
- mcfedr 3y agoThe author shows how to-upper and to-lower change according to locale But making it clear which glyph to use is also a key feature!
- JonChesterfield 3y agoWell C is locale dependent. And one does not break backwards compatibility with C for fear of badness. So naturally Unicode must be locale dependent too.
- layer8 3y agoThis is pretty good. One thing I would add is to mention that Unicode defines algorithms for bidirectional text, collation (sorting order), line breaking and other text segmentation (words and sentences, besides grapheme clusters). The main point here is to know that there are specifications one should take into account when topics like that come up, instead of just inventing your own algorithm.
- dathinab 3y agoThe author seem to hate people which concentration issues and/or various visual sicknesses. That coloration tools shows the moving mouse coursers of other participants even if they aren't needed/wanted is already pretty bad, why bring it to a website?
- wffurr 3y agoThis seems like good feedback but it could really be phrased more constructively. I doubt the author “hates” any such thing and you know it too. “Didn’t design with such in mind”, sure. You can do better.
- dathinab 3y agoyes I should have highlighted that it is satire through it also wasn't meant to be constructive critique
- tr888 3y agoWhat on EARTH is that mouse cursor thing all about? Why would you even bother writing this, then making it impossible to read properly?
- eerikkivistik 3y agoI stopped in the middle of reading the post just for this. It was so distracting I was unable to focus on the text. It's a fun gimmick, but the result is that someone who wanted to read the post, stopped in the middle.
- yrro 3y agoI got a good laugh out of it.
- oliwarner 3y agoIt's tracking every visitors' cursor and sharing it with every other visitor. Why would a frontend developer demonstrate their ability to do frontend programming on their personal, not altogether super-serious blog? I meant that rhetorically but it's a flex. I agree, not the best design in the world if you're catering for particular needs, but simple and fun enough. You should check out dark mode. In that vein, I think it's okay if we let people have fun. That might not work for everyone, but why should we let perfect be the worst enemy of fun?
- dathinab 3y ago> Why would because it shows that they don't understand important design aspects while it doesn't really show off their technical skills because it could be some plugin or copy pasted code, only someone who looks at the code would know better. But if someone care enough about you to look at your code you don't need to show of that skill on you normal web-site and can have some separate tech demo. > okay if we let people have fun yes people having fun is always fine especially if you don't care if anyone ever reads your blog or looks at it for whatever reason (e.g. hiring) but the moment you want people to look at it for whatever reason then there is tension i.e. people don't get hired to have fun and if you want others to read you blog you probably shouldn't assault them with constant distractions
- dathinab 3y ago> The only modern language that gets it right is Swift: I disagree. What is the "right" things is use-case dependent. For UI it's glyph bases, kinda, more precise some good enough abstraction over render width. For which glyphs are not always good enough but also the best you can get without adding a ton of complexity. But for pretty much every other use-case you want storage byte size. I mean in the UI you care about the length of a string because there is limited width to render a strings. But everywhere else you care about it because of (memory) resource limitations and costs in various ways. Weather that is for bandwidth cost, storage cost, number of network packages, efficient index-ability, etc. etc. In rare cases being able to type it, but then it's often us-ascii only, too.
- hot_gril 3y agoSwift made an effort to handle grapheme clusters but severely over-complicated strings by exposing performance details to users. Look at the complex SO answers to what should be simple questions, like finding a substring: https://news.ycombinator.com/item?id=32325511 https://news.ycombinator.com/item?id=32325511 , many of which changed several times between Swift versions I was working on an app in Swift that needed full emoji support once. Team ended up writing our own string lib that stores things as an array of single-character Swift strings.
- marcellus23 3y ago> many of which changed several times between Swift versions This was true while Swift was developing but it's been stable now for several years. At some point that complaint is no longer valid.
- hot_gril 3y agoYou still see all the answers from old versions sitting around, often at the top. Part of it is because of how often they changed such fundamental things. String length changed 3 times. Every other language figured these things out before the initial non-beta release.
- badcppdev 3y agoJust a nitpick because the page says: "Unicode is a standard that aims to unify all human languages, both past and present, and make them work with computers." but of course unicode is only relevant to written languages as opposed to spoken languages (and signed languages) I wish that was the only thing wrong with that page
- jordanrobinson 3y agoAnyone know what the story is behind the "Weird Emoji" around 140000 on the map?
- Findecanor 3y agoThe E0000-E007F block is the "Tags" block, which is used for flag emojis. But there is not a code for each flag. Instead there is a code for each ASCII character. A flag sequence is formed from U+1F3F4 (Black Flag), followed by at least two tags that form a country/region code, and then U+E007F (End tag). So, yes this is weird, because the emoji is dependent on the decoder. It was made this way to keep Unicode independent of geopolitics. Read more: <https://en.wikipedia.org/wiki/Tags_(Unicode_block) https://en.wikipedia.org/wiki/Tags_(Unicode_block)>
- RugnirViking 3y agohttps://www.unicode.org/charts/PDF/UE0000.pdf https://www.unicode.org/charts/PDF/UE0000.pdf
- layer8 3y agohttps://archive.ph/LtKk0 https://archive.ph/LtKk0
- neonate 3y agohttp://web.archive.org/web/20231002163213/https://tonsky.me/blog/unicode/ http://web.archive.org/web/20231002163213/https://tonsky.me/...
- Karellen 3y ago> The simplest possible encoding for Unicode is UTF-32. It simply stores code points as 32-bit integers. Skipping over UTF-32-BE and UTF-32-LE there... (I mean, it might not be an issue if it's just being used as an internal representation, but still)
- deleted 3y ago[deleted]
- zzzeek 3y agoI wondered about how to do simple text centering / spacing justification given graphemes showing string lengths that don't match up human-perceived characters, like in 'Café' (python len('Café') returns 5, even though we see four letters). Found this! good to know about. https://pypi.org/project/grapheme/ https://pypi.org/project/grapheme/ "A Python package for working with user perceived characters. " (apparently the article talks about this however the blog post is largely unreadable due to dozens of animated arrow pointers jumping all over the screen)
- ggcampinho 3y agoElixir also gets the length correctly, not only Swift.
- gumby 3y agoThis is quite a good write up. An answer to one of the author's questions: > Why does the fi ligature even have its own code point? No idea. On of the principles of Unicode is round trip compatibility. That is you should be able to read in a file encoded with some obsolete coding system and write it out again properly. Maybe frob it a bit with your unicode-based tools first. This is a good principle, though less useful today. So the fi ligature was in a legacy encoding system and thus must be in Unicode. That's also why things like digits with a circle around them exist: they were in some old Japanese character set. Nowadays we might compose them with some zwj or even just leave them to some higher level formatting (my preference).
- gwervc 3y agoThe circled digits as code points are very nice to have precisely because they are available in applications that don't support them otherwise... which is actually most of the software I can think of (Notepad, Apple Notes, chat applications, most websites, etc).
- swores 3y agoCan you write them with iOS keyboard? Or when you say Apple Notes and chat apps you just mean from desktop? Edit ①: seems the answer is not with the default iOS keyboard, but possible to paste it and perhaps possible with a third party keyboard that I'm not keen on trying (unless I hear of a keyboard that's both genuinely useful / better than default, and that doesn't send keystrokes to the developer - though I can't remember if the latter is even a risk on iOS, better go search about that next..)
- masklinn 3y agoYou can copy/paste them from a character board, a dedicated website, or even the wiki.
- d11z 3y agoSpeaking of third party keyboards, I’m still upset about what happened to Nintype[0]. I’ve never ever been able to type faster on mobile than with it’s intuitive hybrid input style of sliding and tapping, paired with AI that was actually good. It used to be quite performant, fully customizable, and it worked beautifully as a replacement for default on jailbroken iOS. Today, it’s buggy $5 abandonware that only makes me sad when I am reminded of it. EDIT: Here[1] is a blog post that claims it's still the best keyboard in 2023. I actually might give it another shot... Not holding my breath though. *EDIT: Looks like another dedicated fan has actually taken it upon themself to revive the project, under the new name Keyboard71[2]. [0] https://apps.apple.com/us/app/nintype/id796959534 https://apps.apple.com/us/app/nintype/id796959534 [1] https://maxleiter.com/blog/nintype https://maxleiter.com/blog/nintype [2] https://www.reddit.com/r/keyboard71/ https://www.reddit.com/r/keyboard71/
- oefrha 3y ago> many Chinese, Japanese, and Korean logograms that are written very differently get assigned the same code point This leads to absolutely horrendous rendering of Chinese filenames in Windows if the system locale isn’t Chinese. The characters seem to be rendered in some variant of MS Gothic and it’s very obviously a mix of Chinese and Japanese glyphs (of somewhat different sizes and/or stroke widths IIRC). I think the Chinese locale avoids the issue by using Microsoft YaHei UI.
- samatman 3y agoPlease don't refer to codepoints as characters. Some are, some are not, it isn't a useful or informative approximation, it's just wrong. Unicode is a table which assigns unique numbers to different codepoints, most of which are characters. ZWJ is not a character at all, and extended grapheme clusters made of several codepoints are.
- skitter 3y ago'Character' doesn't have a single meaning. ZWJ is a character according to definitions (2) and (3) in https://unicode.org/glossary/#character https://unicode.org/glossary/#character
- justrealist 3y agoI don't want to be too full of myself here, but I'm a very skilled and highly paid backend software engineer who knows roughly nothing about unicode (I google what I need when a file seems f'd up), and it's never been a problem for me. I'm sure the article is good but the title is nonsense.
- bigstrat2003 3y agoThe title is definitely nonsense. The reality is that for most people, they will never need to know the gritty details of how to encode or decode UTF-8. The article is interesting, but I was pretty put off with how the author led with such a hyperbolic (and untrue) claim.
- bagasme 3y agoThe article doesn't mention how to resolve string manipulation problem involving locales.
- m3kw9 3y agoUnicode looks like a big over engineered standard that had 50 hands trying to put their mark in
- pornel 3y agoThat's because Unicode chose to be a superset of all other encodings, so they've brought everyone else's complexity and tech debt.
- ssokolow 3y agoTechnically, a superset would have to somehow Schrödinger's cat around \ in latin1 and ¥ in Shift-JIS being the same codepoint. Unicode just took it upon themselves to reliably round-trip legacy text... thus the precomposed forms. Most of the other complexity and technical debt is in the writing systems themselves.
- ebiester 3y agoIt looks like that because Unicode is trying to solve a problem that everyone thinks is easy until they uncover the true extent of encoding human languages.
- eviks 3y agoHow does this explain surrogate pairs?
- jfultz 3y agoSurrogate pairs were new to Unicode 2.0. Unicode 1.0 didn't anticipate the need for more than 65,536 code points (who would ever need more?); the main perceived threat to that limit having been resolved by Han unification.
- eviks 3y agoOk, but that doesn't answer the question; it's more of an indication that those design(at)s didn't uncover "the true extent" until years later
- JonChesterfield 3y agoPrior to this article, I knew graphemes were a thing and that proper unicode software is supposed to count those instead of bytes or code points. I didn't know that unicode changes the definition of grapheme in backwards incompatible fashion annually, so software which works by grapheme count is probably inconsistent with other software using a different version of the standard anyway. I'm therefore going to continue counting bytes. And comparing by memcmp. If the bytes look like unicode to some reader, fine. Opaque string as far as my software is concerned.
- slimsag 3y agoTwo Unicode strings can be visually and semantically identical, but not byte-equal.
- dundarious 3y agoThe point is that a byte focus will often frustrate users. e.g., a TUI with columns will have to truncate "long" strings in each column, and that truncation and column-separator arrangement really should be grapheme aware. e.g., a string search (for a name, let's say) should find Noël regardless of whether the user input ë via composing characters or the pre-composed version.
- tonsky 3y agoGood luck https://mastodon.online/@alexeyten@mas.to/111166351426290784 https://mastodon.online/@alexeyten@mas.to/111166351426290784
- PeterisP 3y agoComparing by memcmp will result in false negatives unless you can ensure that all incoming text gets normalized to a particular canonical form.
- jandrese 3y agoEven then there are minefields in comparing text, especially case insensitive matching and supporting CKJ.
- 3y ago
- amelius 3y agoCan we please get a standard that describes how emoji are supposed to look? Now they look different on every platform and many subtleties are lost in translation.
- JohnFen 3y agoYeah, this problem has led me to avoid using emojis. I can't be sure that the meaning I was intending is the one being depicted by the recipients machine. It's probably a good thing, though.