18 ms·
How UTF-8 Works
- who-shot-jr 5y agoFantastic! Very well explained.
- SethMLarson 5y agoThanks for the kind comment :)
- BitwiseFool 5y agoI feel the same way as the GP, great work. I also appreciate how clean and easy to read the diagrams are.
- dspillett 5y agoNot sure if the issue is with Chrome or my local config generally (bog standard Windows, nothing fancy), but the us-flag example doesn't render as intended. It shows as "US" with the components in the next step being "U" and "S" (not the ASCII characters U & S, the encoding is as intended but those characters are being given in place of the intended). Displays as I assume intended in Firefox on the same machine: American flag emoji then when broken down in the next step U-in-a-box & S-in-a-box. The other examples seem fine in Chrome. Take care when using relatively new additions to the Unicode emoji-set, test to make sure your intentions are correctly displayed in all the brower's you might expect your audience to be using.
- SethMLarson 5y agoYeah, there's not much I can do there unfortunately (since I'm using SVG with the actual U and S emojis to show the flag). I can't comment on whether it's your config or not, but I've tested the SVGs on iOS and Firefox/Chrome on desktop to make sure they rendered nicely for most people. Sorry you aren't getting a great experience there. Here's how it's rendering for me on Firefox: https://pasteboard.co/rjLtqANVQUIJ.png https://pasteboard.co/rjLtqANVQUIJ.png
- xurukefi 5y agoFor me it also renders like this on Chrome/Windows: https://i.imgur.com/HCJTpfA.png https://i.imgur.com/HCJTpfA.png Really nice diagrams nevertheless
- dspillett 5y agoYep, that is what I get in FF to. And Chrome on Android, though with a rendering difference on the not-plain-ol'-U and not-plain-ol'-S (in both cases, blue letters rather than white with a blue background).
- andylynch 5y agoThey aren't new (2010) - this is a Windows thing - speculation is it's a policy decision to avoid awkward conversations with various governments (presumably large customers) about TW , PS and others -- see long discussion here for instance https://answers.microsoft.com/en-us/windows/forum/all/flag-emoji/85b163bc-786a-4918-9042-763ccf4b6c05 https://answers.microsoft.com/en-us/windows/forum/all/flag-e...
- mark-r 5y agoUTF-8 is one of the most brilliant things I've ever seen. I only wish it had been invented and caught on before so many influential bodies started using UCS-2 instead.
- SethMLarson 5y ago100% agree, it's really rare that there's a ~blanket solution to a whole class of problems. "Just use UTF-8!"
- josephg 5y agoAbsolutely. At least it’s well supported now in very old languages (like C) and very new languages (like Rust). But Java, Javascript, C# and others will probably be stuck using UCS-2 forever.
- stewx 5y agoWhat is stopping people from encoding their Java, JS, and C# files in UTF-8?
- maskros 5y agoNothing, but Java's "char" type is always going to be 16-bit.
- josephg 5y agoYep. In javascript (and Java and C# from memory) the String.length property is based on the encoding length in UTF16. It’s essentially useless. I don’t know if I’ve ever seen a valid use for the javascript String.length field in a program which handles Unicode correctly. There’s 3 valid (and useful) ways to measure a string depending on context: - Number of Unicode characters (useful in collaborative editing) - Byte length when encoded (these days usually in utf8) - and the number of rendered grapheme clusters All of these measures are identical in ASCII text - which is an endless source of bugs. Sadly these languages give you a deceptively useless .length property and make you go fishing when you want to make your code correct.
- jjice 5y agoFun fact: Ken Thompson and Rob Pike of Unix, Plan 9, Go, and other fame had a heavy influence on the standard while working on Plan 9. To quote Wikipedia: > Thompson's design was outlined on September 2, 1992, on a placemat in a New Jersey diner with Rob Pike. If that isn't a classic story of an international standard's creation/impactful update, then I don't know what is. https://en.wikipedia.org/wiki/UTF-8#FSS-UTF https://en.wikipedia.org/wiki/UTF-8#FSS-UTF
- SethMLarson 5y agoI knew that Ken Thompson had an influence but wasn't aware of Rob Pike, what a great fact! Thanks for sharing this :)
- ChrisSD 5y agoFor whatever it's worth Rob Pike seems to credit Ken Thompson for the invention, though they both worked together to make it the encoding used by Plan 9 and to advocate for its use more widely.
- nayuki 5y agoExcellent presentation! One improvement to consider is that many usages of "code point" should be "Unicode scalar value" instead. Basically, you don't want to use UTF-8 to encode UTF-16 surrogate code points (which are not scalar values). Fun fact, UTF-8's prefix scheme can cover up to 31 payload bits. See https://en.wikipedia.org/wiki/UTF-8#FSS-UTF https://en.wikipedia.org/wiki/UTF-8#FSS-UTF , section "FSS-UTF (1992) / UTF-8 (1993)". A manifesto that was much more important ~15 years ago when UTF-8 hadn't completely won yet: https://utf8everywhere.org/ https://utf8everywhere.org/
- masklinn 5y ago> Fun fact, UTF-8's prefix scheme can cover up to 31 payload bits. It’d probably be more correct to say that it was originally defined to cover 31 payload bits: you can easily complete the first byte to get 7 and 8 byte sequences (35 and 41 bits payloads). Alternatively, you could save the 11111111 leading byte to flag the following bytes as counts (5 bits each since you’d need a flag bit to indicate whether this was the last), then add the actual payload afterwards, this would give you an infinite-size payload, though it would make the payload size dynamic and streamed (where currently you can get the entire USV in two fetches, as the first byte tells you exactly how many continuation bytes you need).
- SethMLarson 5y agoYeah the current definition is restricted to 4 octets in RFC 3629. Really interesting to see the history of ranges UTF-8 was able to cover.
- CountSessine 5y agoBasically, you don't want to use UTF-8 to encode UTF-16 surrogate code points The awful truth is that there is such a beast. UTF-8 wrapper with UTF-16 surrogate pairs. https://en.wikipedia.org/wiki/CESU-8 https://en.wikipedia.org/wiki/CESU-8
- nayuki 5y agoIs CESU-8 a synonym of WTF-8? https://en.wikipedia.org/wiki/UTF-8#WTF-8 https://en.wikipedia.org/wiki/UTF-8#WTF-8 ; https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/
- brian_rak 5y agoThis was presented well. A follow up for unicode might be in order!
- SethMLarson 5y agoGlad you enjoyed! Unicode and how it interacts with other aspects of computers (IDNA, NFKC, grapheme clusters, etc) is some of the spaces I want to explore more.
- jsrcout 5y agoThis may be the first explanation of Unicode representation that I can actually follow. Great work.
- SethMLarson 5y agoWow, thank you for the kind words. You've made my morning!!
- filleokus 5y agoRecently I learned about UTF-16 when doing some stuff with PowerShell on Windows. Parallel with my annoyance with Microsoft, I realized how long it’s been since I encountered any kind of text encoding drama. As a regular typer of åäö, many hours of my youth was spent on configuring shells, terminal emulators, and IRC clients to use compatible encodings. The wide adoption of UTF-8 has been truly awesome. Let’s just hope it’s another 15-20 years until I have to deal with UTF-16 again…
- ChrisSD 5y agoThere are many reasons why UTF-8 is a better encoding but UTF-16 does at least have the benefit of being simpler. Every scalar value is either encoded as a single unit or a pair of units (leading surrogate + trailing surrogate). However, Powershell (or more often the host console) has a lot of issues with handling Unicode. This has been improving in recent years but it's still a work in progress.
- tialaramex 5y agoUTF-16 only makes sense if you were sure UCS-2 would be fine, and then oops, Unicode is going to be more than 16-bits and so UCS-2 won't work and you need to somehow cope anyway. It makes zero sense to adopt this in greenfield projects today, whereas Java and Windows, which had bought into UCS-2 back in the early-mid 1990s, needed UTF-16 or else they would need to throw all their 16-bit text APIs away and start over. UTF-32 / UCS-4 is fine but feels very bloated especially if a lot of your text data is more or less ASCII, which if it's not literally human text it usually will be, and feels a bit bloated even on a good day (it's always wasting 11-bits per character!) UTF-8 is a little more complicated to handle than UTF-16 and certainly than UTF-32 but it's nice and compact, it's pretty ASCII compatible (lots of tools that work with ASCII also work fine with UTF-8 unless you insist on adding a spurious UTF-8 "byte order mark" to the front of text) and so it was a huge success once it was designed.
- ChrisSD 5y agoAs I said, there are many reasons UTF-8 is a better encoding. And indeed compact, backwards compatible, encoding of ASCII is one of them.
- daenz 5y agoGreat explanation. The only part that tripped me up was in determining the number of octets to represent the codepoint. From the post: >From the previous diagram the value 0x1F602 falls in the range for a 4 octets header (between 0x10000 and 0x10FFFF) Using the diagram in the post would be a crutch to rely on. It seems easier to remember the maximum number of "data" bits that each octet layout can support (7, 11, 16, 21). Then by knowing that 0x1F602 maps to 11111011000000010, which is 17 bits, you know it must fit into the 4-octet layout, which can hold 21 bits.
- bumblebritches5 5y ago
- mananaysiempre 5y agoAs the continuation bytes always bear the payload in the low 6 bits, Connor Lane Smith suggests writing them out in octal[1]. Though that 3 octets of UTF-8 precisely cover the BMP is also quite convenient and easy to remember (but perhaps don’t use that like MySQL did[2]?..). [1] http://www.lubutu.com/soso/write-out-unicode-in-octal http://www.lubutu.com/soso/write-out-unicode-in-octal [2] https://mathiasbynens.be/notes/mysql-utf8mb4 https://mathiasbynens.be/notes/mysql-utf8mb4
- bussyfumes 5y agoBTW here’s a surprise I had to learn at some point: strings in JS are UTF-16. Keep that in mind if you want to use the console to follow this great article, you’ll get the surrogate pair for the emoji instead.
- zaik 5y agoThose diagrams look really good. How were they made?
- jeremieb 5y agoThe author mentions at the end of the article that he spent a lot of time on https://www.diagrams.net/ https://www.diagrams.net/. :)
- jokoon 5y agoI wonder how large must a font be to display all UTF8 characters... I'm also waiting for new emojis, they recently added more and more that can be used as icons, which is simpler than integrating PNG or SVG icons.
- dspillett 5y agoTake care using recently added Unicode entries, unless you have some control of your user-base and when they update or are providing a custom font that you know has those items represented. You could be giving out broken-looking UI to many if their setup does not interpret the newly assigned codes correctly.
- banana_giraffe 5y agoOpentype makes this impossible. A glyph has an index of a UINT16, so you can't fit all of the ~143k Unicode characters. There are some attempts at font families to cover the majority of characters. Like Noto ( https://fonts.google.com/noto/fonts https://fonts.google.com/noto/fonts ), broken out into different fonts for different regions. Or, Unifont's ( http://www.unifoundry.com/ http://www.unifoundry.com/ ) goal of gathering the first 65536 code points in one font, though it leaves a lot to be desired if you actually use it as a font.
- pierrebai 5y agoI never understood why ITF-8 did not use the much simpler encoding of: - 0xxxxxxx -> 7 bits, ASCII compatible (same as UTF-8) - 10xxxxxx -> 6 bits, more bits to come - 11xxxxxx -> final 6 bits. It has multiple benefits: - It encodes more bits per octet: 7, 12, 18, 24 vs 7, 11, 16, 21 for UTF-8 - It is easily extensible for more bits. - Such extra bits extension is backward compatible for reasonable implementations. The last point is key: UTF-8 would need to invent a new prefix to go beyond 21 bits. Old software would not know the new prefix and what to do with it. With the simpler scheme, they could potentially work out of the box up to at least 30 bits (that's a billion code points, much more than the mere million of 21 bits). The
- LegionMammal978 5y agoThe problem is that UTF-8 has the ability to detect and reject partial characters at the start of the string; this encoding would silently produce an incorrect character. Also, UTF-8 is easily extensible already: the bit patterns 111110xx, 1111110x, and 11111110 are only disallowed for compatibility with UTF-16's limits.
- pierrebai 5y agoHow often are stream truncated at the start? In my career, I've seen plenty of end truncation, but start truncation never happens. Or, to be more precise, it only happens if previous decoding is already borked. If a previous decoding read too much data, then even UTF-8 is borked. You could be decoding UTF-8 from the bits of any follow-up data. Even for pure text data, if a previous field was over-read (the only plausible way to have start-truncation), then you probably are decoding incorrect data from then on. IOW, this upside is both ludicrously improbable and much more damning to the decoding than simply be able to skip a character.
- maxdamantus 5y agoIt happens all the time when you're working with substrings. Under your encoding, "é" would be a substring of "ჩ": 10000011 11101001 U+E9 "é" 10000001 10000011 11101001 U+10E9 "ჩ" This would make it incompatible with many existing processes that already handle text, which was one of the goals of UTF-8.
- ctxc 5y agoSuch clean presentation, refreshing.
- riwsky 5y agoHow UTF-8 works? “pretty well, all things considered”
- karsinkk 5y agoThe following article is one of my favorite primers on Character sets/Unicode : https://www.joelonsoftware.com/2003/10/08/the-absolute-minim https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
- DannyB2 5y agoThere is an error in the first example under Giant Reference Card. The bytes come out as: 0xF0 0x9F 0x87 0xBA 0xF0 0x9F 0x87 0xBA but the bits directly above them all of the bit pattern: 010 10111
- SethMLarson 5y agoGreat eye! I'll fix this and push it out.
- jvolkman 5y agoRob Pike wrote up his version of its inception almost 20 years ago. The history of UTF-8 as told by Rob Pike (2003): http://doc.cat-v.org/bell_labs/utf-8_history http://doc.cat-v.org/bell_labs/utf-8_history Recent HN discussion: https://news.ycombinator.com/item?id=26735958 https://news.ycombinator.com/item?id=26735958
- bhawks 5y agoUtf8 is one of the most momentous and under appreciated / relatively unknown achievements in software. A sketch on a diner placemat has lead to every person in the world being able to communicate written language digitally using a common software stack. Thanks to Ken Thompson and Rob Pike we have avoided the deeply siloed and incompatible world that code pages, wide chars and other insufficient encoding schemes were guiding us towards.
- ahelwer 5y agoIt really is wonderful. I was forced to wrap my head around it in the past year while writing a tree-sitter grammar for a language that supports Unicode. Calculating column position gets a whole lot trickier when the preceding codepoints are of variable byte-width! It's one of those rabbit holes where you can see people whose entire career is wrapped up in incredibly tiny details like what number maps to what symbol - and it can get real political!
- deleted 5y ago[deleted]
- inglor_cz 5y agoAs a young Czech programming acolyte in the late 1990s, I had to cope with several competing 8-bit encodings. It was a pure nightmare. Long live UTF-8. Finally I can write any Central European name without mutilating it.
- cryptonector 5y agoAnd stayed ASCII-compatible. And did not have to go to wide chars. And it does not suck. And it resynchronizes. And...
- wolverine876 5y ago> And did not have to go to wide chars. What do you mean by that? Unicode does have double-wide characters and, I discovered recently, some other characters called 'wide' that at least are wider than one column but smaller than two in at least some monospaced fonts. Try: * Small Hyphen-minus (U+FE63) "﹣": Seems to be >1 and <2 columns in at least some monospaced fonts. * Fullwidth Hyphen-minus (U+FF0D) "-": ditto
- RoddaWallPro 5y agoI spent 2 hours last Friday trying to wrap my head around what UTF-8 was (https://www.joelonsoftware.com/2003/10/08/the-absolute-minim https://www.joelonsoftware.com/2003/10/08/the-absolute-minim is great, but doesn't explain the inner workings like this does) and completely failed, could not understand it. This made it super easy to grok, thank you!
- nabla9 5y ago>NOTE: You can always find a character boundary from an arbitrary point in a stream of octets by moving left an octet each time the current octet starts with the bit prefix 10 which indicates a tail octet. At most you'll have to move left 3 octets to find the nearest header octet. This is incorrect. You can only find boundaries between code points this way. Until your you learn that not all "user perceived characters" (grapheme clusters) can be expressed as single code point Unicode seems cool. These UTF-8 explanations explain the encoding but leave out this unfortunate detail. Author might not even know this because they deal with subset of Unicode in their life. If you want to split text between two user perceived characters, not between them, this tutorial does not help. Unicode encodings are is great if you want to handle subset of languages and characters, if you want to be complete, it's a mess.
- SethMLarson 5y agoYou're right, that should read "codepoint boundary" not "character boundary". I can fix that. I do briefly mention grapheme clusters near the end, didn't want to introduce them as this article was more about the encoding mechanism itself. Maybe a future article after more research :)
- nabla9 5y agoPlease do. You have the best visualizations of UTF-8 I have seen so far. Usually people write just the UTF-8 encoding part, then don't mention the rest of the Unicode, because it's clearly not as good and simple.
- YaBomm 5y ago
- Simplicitas 5y agoI still wanna know in WHICH Jersey diner it was invented in! :-)
- burtekd 5y agoI love Tom Scott's explanation of Unicode: https://www.youtube.com/watch?v=MijmeoH9LT4 https://www.youtube.com/watch?v=MijmeoH9LT4
- satysin 5y agoThis is without question one of the best short technical presentations I've seen. To the author hats off to a masterful job.
- geokon 5y agoIs there any standard system where each byte/word maps to one character/grapheme? I feel there is a general sentiment that not being able to jump to the Nth character is programmatically.. irritating and disappointing. I'm sure such a system wouldn't support some languages - but in the words of Lord Farquaad "That's a sacrifice I'm willing to make". Most of the world's languages would do just fine and it'd make sense to exclude right to left arabic ligatured text in, for instance, your monospaced computer code. I'm guessing you could extract a subset of UTF-8 - but has anyone done anything like that?
- deleted 5y ago[deleted]
- eyelidlessness 5y agoYou’re basically describing the various character encodings which preceded adoption of Unicode. No, UTF-8 doesn’t support that. That’s the sacrifice it was willing to make, to eliminate the need to distinguish between character encodings: a string isn’t a sequence of bytes. Even if you don’t appreciate that convenience (and the highly standardized, easily abstracted rules of combined bytes), maybe you’ll appreciate that the complexity of supporting multilingual users would mean multipart data, arbitrarily fractured. As in a single text message using multiple character sets could be dozens of payload boundaries. Does that really sound easier or more painless than UTF-8’s variable-length code points?
- geokon 5y agoFrom what I understood of the link there is the "grapheme" which is an intermediary representation between being UTF-8 and being drawn on the screen. Like all possible glyphs have been prerendered. It's not got some simple graphics engine to draw those on the fly. Couldn't you directly provide an index in the grapheme "cache" ?
- eyelidlessness 5y agoThe index is larger than you think.
- alblue 5y agoIf you’re more into watching a presentation, I recorded “A Brief History of Unicode” last year, And there’s a YouTube recording of it as well as the slides: https://speakerdeck.com/alblue/a-brief-history-of-unicode-4524a734-aac3-4ce9-8c4a-6f4ada04f464 https://speakerdeck.com/alblue/a-brief-history-of-unicode-45... https://youtu.be/NN3g4JbbjTE https://youtu.be/NN3g4JbbjTE
- devstein 5y agoGreat post and intuitive visuals! I recently had to rack my brain around UTF-8 encoding and decoding when building the Unicode ETH Project (https://github.com/devstein/unicode-eth https://github.com/devstein/unicode-eth) and this post would have been very useful
- loco5niner 5y agoExcellent article, really helped me learn. I'd like to add a correction. The binary/ascii/utf-8 value of 'a' (hex 0x61) is not 01010111, but instead 01100001. This is used incorrectly in both the Giant reference card, and in the "ascii encoding" diagram above it.