11 ms·
The UTF-8-Everywhere Manifesto
- tommi 14y agoThat collection of best practices can hardly be considered as "UTF-8 Everywhere Manifesto" as it focuses on Windows and C++. It's good, but I'd rather see more manifesto like document for all cases on a domain like that.
- archangel_one 14y agoI suspect this is mainly because Windows C++ programmers are the largest group that they feel need convincing. Which isn't totally their fault, Microsoft haven't done well by them by not offering good support for UTF-8; you can convert to/from it using WideCharToMultiByte but that's pretty low level, and higher-level APIs like CString will cheerfully munge UTF-8 strings for you. They also tend to conflate Unicode and UTF-16 which again doesn't help less experienced programmers realise that there might be alternatives. I've been through the Windows Unicode stuff at a previous job, which ended up using mostly UTF-16 with some UTF-8 for interfacing to third party libraries and for files which needed to be backward compatible to ASCII (plus significant space savings, which I fought hard for). I think I prefer that approach though, since after the (difficult) conversion you didn't need to worry about encodings in 99% of the code. By their rules you'd gain significant complexity by transforming all over the place in any non-trivial GUI code.
- deleted 14y ago[deleted]
- raverbashing 14y agoDisagree "UTF-16 is the worst of both worlds—variable length and too wide" Really, the author tries to convince the reader, but it's not that clean cut. One of the advantages of UTF-16 is knowing right away it's UTF-16 as opposed to deciding if it's UTF-8/ASCII/other encoding. Sure, for transmission it's a waste of space (still, text for today's computer capabilities is a non issue even if using UTF-32) "It's not fixed width" But for most text, it is. Sure, you can do UTF-32 and it may not be a bad idea (today) Yes, Windows has to deal with several complications and with backwards compatibility, so it's a bag of hurt. Still, they went the right way (internally, it's unicode, period.) "in plain Windows edit control (until Vista), it takes two backspaces to delete a character which takes 4 bytes in UTF-16" If I'm not mistaken this is by design. The 4 byte characters is usually typed as a combination of characters, so if you want to change the last part of the combination you jut type one backspace.
- fauigerzigerk 14y agostill, text for today's computer capabilities is a non issue even if using UTF-32 That obviously depends entirely on what kind of application we're talking about. Keeping large amounts of text data in memory as efficiently as possible is one of my greatest concerns. Many people are processing lots of text nowadays, more than ever before. "It's not fixed width" But for most text, it is. True, so ignoring it means that your code will be correct ... most of the time.
- raverbashing 14y ago"Keeping large amounts of text data in memory as efficiently as possible is one of my greatest concerns" Depends on what you consider efficiency. If it's size, sure, store it using UTF-8. But if you're worried about speed, then UTF-16 or 32 may be the way to go, since you're dealing with data that fits a CPU register. For example, on ARM comparing one byte is much more work than comparing one 32-bit value. "True, so ignoring it means that your code will be correct ... most of the time." No, not going to ignore it! But on UTF-16 more code points match the UTF-16 encoding of it (easier for debugging)
- daliusd 14y ago"One of the advantages of UTF-16 is knowing right away it's UTF-16 as opposed to deciding if it's UTF-8/ASCII/other encoding." It is actually not that simple. By using UTF-16 you already have at least two problems: 1. You should know if byte order is big endian or little endian. 2. You should know if your API supports whole unicode set or only 65536 symbols. E.g. Windows API. Do you know answer? What will happen if your user wants to abuse your system by using symbols outside those 65536.
- paulsutter 14y agoI think the author's basic point is that if we standardize on utf-8, that "8 bit anxiety" goes away. I did my first programming assignment on punched cards, so I probably have permanent ASCII/EBCDIC brain damage. However, this article decisively convinces even me that utf-8 wins and other encodings represent fail.
- 14y ago
- breck 14y agoHow could we avoid acronyms like 'utf-8'? We can do better than that. Unicode8?
- _ak 14y agoWhy should we?
- soult 14y agoSo we should call it "Unicode-8" to further confuse people about the difference between "Unicode" and "UTF-8"?
- paulsutter 14y agoJust use the term "string" to refer to utf-8, and the term "data in nonstandard encoding X" to refer to other encodings. In the article he puts in in terms of std::string, but more generally I think this is what he means.
- adamtj 14y agoYou're confusing things. Strings cannot be utf-8 any more than you can be your signature. "strings" are abstract data structures. They are lists of characters. Not bytes, not integers, but characters. Often, we use the Unicode character set as the set of allowable characters. There are other character sets. Internally, strings often represent characters as integers. When using the Unicode character set, strings then use the Unicode encoding to integers (a table mapping characters to unique numbers). Sometimes we use other character sets and encodings. Unfortunately, integers are abstract. You can't store them in a file or transmit them over a network until you pick a concrete representation as bytes. How many bits per integer? Big or little endian? Etc. That's where UTF-8 comes into play. UTF-8 is a merely a compressed data format used to represent a sequence of integers as a sequence of bytes - a way that happens to have some properties that make it convenient for representing strings. UTF-8 is not Unicode. UTF-8 can also be used for other types of numerical data. As a silly example, suppose you had a list of ages of houses. Many houses are less than 100 years old. A few are more than 300 years old. An efficient serialization of that data would be to represent the ages as integers and then utf-8 encode your list of integers. Some true statements: A character set is a set of characters. Characters are not integers or bytes. A mapping from characters to integers is an encoding. Unicode is a standard that defines a character set and an encoding to integers. Mapping integers to bytes is confusingly also called encoding. UTF-8 is an encoding from integers to bytes. Unicode defines a set of characters and an encoding of characters to integers. UTF-8 is an encoding of integers to bytes. UTF-8 is not Unicode.
- alecco 14y agoASCII and UTF-8 are too US centric. That's why adoption in places like China is so low. Also, if there's variable length encoding why can't we just do a proper way and improve size for the same computational cost?
- paulsutter 14y agoThe author makes a compelling case for UTF-8 in Asian languages. I'd love to hear any specific counter-arguments.
- maaku 14y agoNo he doesn't, he dismisses it out of hand by choosing an example that is uniquely suited to minimize the advantages of UTF-16 for non-Roman scripts. Precious little of an HTML document is actually textual content.
- ori_b 14y agoI think that's the point; In the real world, any significant chunk of non-roman text is embedded in far more roman text, or in terms of internals of programs, is generally dwarfed by the size of other data structures. More or less, I'd say that storage size of text usually doesn't matter.
- paulsutter 14y agoPrecious little of the data stored in the world is textual content. It's hard to believe that choice of encoding standards materially impacts RAM or disk budgets. Most of his points are not related to encoding size but to simplicity and standardization. I find those reasons to be very compelling. I'm asking for clear counterarguments because I concede that my ASCII background could predispose me to UTF-8. Please be a little clearer, I really do want to know.
- pjscott 14y agoHTML documents are hardly unusual examples. Also, look at the other column in that table, where he stripped out the HTML tags and looked only at the body text: UTF-16 was somewhat smaller, and gzipping them made the difference negligible. Does UTF-16 really have such a great advantage for non-Roman writing systems? Or is this motivated more by a disliking for Anglocentrism?
- gwillen 14y agoOk, let me be the first approving top level comment: This document is correct. The author of this document is smart. You should follow this document. As jwz said about backups: "Shut up. I know things. You will listen to me. Do it anyway."
- sopooneo 14y agoCan someone explain to me how UTF-8 is endianness independent? I don't mean that I am arguing the fact, I just don't understand how it is possible. Don't you have to know which order to interpret the bits in each byte? And isn't that endianness?
- evincarofautumn 14y agoNo, that’s not endianness; endianness refers to the ordering of bytes within a multi-byte value—least significant byte first or most significant byte first, generally. The order of octets in a UTF-8 code point is fixed, and because a bit is not an addressable unit of memory, the storage order of bits within an octet is immaterial.
- kijin 14y agoIt's endianness independent in the sense that the order in which you interpret the bytes in each character does not depend on the processor architecture, unlike UTF-16. If your processor interprets the bits in each byte in a different order, that might be a problem, but it's not what we're talking about when we usually talk about the endianness of character encodings. http://en.wikipedia.org/wiki/Endianness http://en.wikipedia.org/wiki/Endianness
- sopooneo 14y agoThank you. That is very good to learn and I looked over the wikipedia article. But as far as byte order, how is that architecture independent? Is it just that utf-8 dictates that the order of the bytes always be the same, so whatever system you're on, you ignore its norm, and interpret bytes in the order utf-8 tells you to?
- ori_b 14y agoutf-8 is a single byte encoding. Reversing the order of a sequence that's one byte long just gives back that one byte.
- 14y ago
- evincarofautumn 14y agoFor those who don’t know it, UTF8-CPP[1] is a good lightweight header-only library for UTF conversions, mostly STL-compatible. [1] http://utfcpp.sourceforge.net/ http://utfcpp.sourceforge.net/
- pcwalton 14y agoSadly, the pervasiveness of JavaScript means that UTF-16 interoperability will be needed as least as long as the Web is alive. JavaScript strings are fundamentally UTF-16. This is why we've tentatively decided to go with UTF-16 in Servo (the experimental browser engine) -- converting to UTF-8 every time text needed to go through the layout engine would kill us in benchmarks. For new APIs in which legacy interoperability isn't needed, I completely approve of this document.
- nknight 14y ago> converting to UTF-8 every time text needed to go through the layout engine would kill us in benchmarks So what? Is your goal to create useful software, or win at worthless benchmarks?
- pcwalton 14y agoWell, we're talking about DOM manipulation performance here. Pages that use DOM manipulation heavily will see a potentially-unacceptable performance loss if text always has to be converted to UTF-8. Is fast DOM manipulation important? Given that the only way for the sole scripting language on the Web to display anything or interact with the user is through DOM manipulation, I think it's worth optimizing every cycle...
- ori_b 14y agohttp://www.utf8everywhere.org/#faq.cvt.perf http://www.utf8everywhere.org/#faq.cvt.perf If the function you're calling with UTF8 is non-trivial, converting a few dozen bytes is unlikely to make a significant difference. Benchmark it, of course, but don't be surprised if you don't need to care. Modifying the DOM is probably going to be non-trivial.
- pcwalton 14y agoSee my comment above; I can construct cases that make this sort of conversion have unacceptable overhead. Would it matter in a real-world setting? I can't say for sure, because nobody I know of has tried making a production-quality UTF-8 web layout engine. But, in my mind, none of the benefits of UTF-8 (memory usage being the main one in a browser [1]) outweigh the performance risks of doing conversion. And the risk is real. [1]: Note that you still need UTF-16 anyway, for interoperability with JavaScript. So using UTF-8 might even lead to worse memory usage, due to the necessity of duplicating strings, than a careful UTF-16-everywhere scheme that takes advantage of string buffer sharing between the JS heap and the layout engine heap would.
- makecheck 14y agoMarkus Kuhn's web page has a lot of useful UTF-8 info and valuable links (e.g. samples of UTF-8 corner cases that people often miss). http://www.cl.cam.ac.uk/~mgk25/unicode.html http://www.cl.cam.ac.uk/~mgk25/unicode.html
- lubutu 14y agoThis is a great resource; it was extremely useful when I was writing a UTF-8 library myself. I found the UTF-8 stress test file is particularly useful to run tests against: http://www.cl.cam.ac.uk/~mgk25/ucs/examples/UTF-8-test.txt http://www.cl.cam.ac.uk/~mgk25/ucs/examples/UTF-8-test.txt
- erichocean 14y agoHmm, TextMate has problems with "5.2 Paired UTF-16 surrogates" in that stress test file. (Yes, I interpreted the file as UTF-8 in TextMate).
- luriel 14y agoYes! I have been meaning to write something like this for years. There is only one thing I would add: Never add a BOM to an UTF-8 file!! It is redundant, useless and breaks all kinds of things by attaching garbage to the start of your files. Edit: Here is the interesting story of how Ken Thompson invented UTF-8: http://doc.cat-v.org/bell_labs/utf-8_history http://doc.cat-v.org/bell_labs/utf-8_history
- makecheck 14y agoThe mark isn't useless; it clearly identifies files as UTF-8 so they can be processed as such immediately. Otherwise a program has to "sniff" several bytes to see if the encoding could be something different, and it may not guess correctly. Also, how can "all kinds of things" break with this mark? If something is reading UTF-8 correctly then it'll be fine with the mark; and if it's not reading UTF-8 correctly then it will screw up a lot more than the mark at the beginning of the file.
- CJefferson 14y agoI there a simple set of rules for people who currently have code which use ASCII, to check for UTF-8 cleanness? In particular, what should I watch out for to make an ASCII parser UTF-8 clean?
- dlubarov 14y agoAll 7-bit ASCII is valid UTF-8, so you're fine as long as you're really using 7-bit ASCII and not latin1 or Windows-1252, etc.
- rogerbinns 14y agoIf you are "parsing" a string then you will have problems unless you specifically make the code deal with unicode code points and not bytes. If you just accept a char* and then pass it on with the contents as is you'll generally be fine (except on Windows).
- makecheck 14y agoIf you're reading something in pieces, like a buffer that fills 256 bytes at a time, you have to be careful. UTF-8 is a multi-byte encoding so the last byte in your buffer may not completely finish a code point. Unlike older code that can just read a bunch of bytes and use them, with multi-byte encodings you have to have a way to deal with "left-overs" until new bytes show up. Fortunately the UTF-8 encoding (e.g. see the Wikipedia page) makes it clear when a byte is the beginning of a new point and it tells you how many intermediate bytes should follow.
- pornel 14y agoUTF-8 is easy to figure out. However, it's only a doorway to Unicode, and Unicode is not simple. If you do anything more than chopping codepoints and passing them around — use a library. There's a lot of technical complexity due to quirks of Unicode and inherent complexity of world's diverse writing systems. Avoid using concept of a "character" as much as you can, as it's fuzzy, e.g. there are combining characters and ligatures (codepoint != character). Be aware that string comparison cannot be done just by comparing codepoints, and there are different levels of "sameness" of Unicode strings coming from different normalisations, e.g. NFC and NFKD. Case-insensitive comparison cannot be done by lowercasing a string: http://www.moserware.com/2008/02/does-your-code-pass-turkey-test.html http://www.moserware.com/2008/02/does-your-code-pass-turkey-... Unfortunately I don't know much about RTL text, and there's lots of traps there too (e.g. there are control characters for controlling text direction).
- natch 14y agoStrings (NSString) on Apple platforms are UTF-16. The Apple platforms are not exactly lagging behind in either multilingual, or text processing. I wonder what this team of three people knows that Apple doesn't? Or is it the other way around, that Apple knows something they don't, and when it comes to shipping products that work in the real world, Apple has figured out how to do it?
- deleted 14y ago[deleted]
- Xuzz 14y agoNSString is decades old, from NeXTSTEP (you can see that in the name: "NS"). While its possible they could change it, the in-memory representation doesn't matter much in this case. When you transfer data out, such as with writeToFile:encoding: or convert it into a char * (often with UTF8String, or one of the C string methods), you are almost always specifying an encoding anyway. And, for most Cocoa apps, that encoding is UTF8.
- gregschlom 14y agoAs the authors explain in the post, many things use UTF-16 internally (Python, Java, C#, etc...). It does not mean that it's the best solution. Have you read the article? They make a good explanation at why UTF-16 is "the worst of both worlds" (wide characters AND variable lenght).
- __david__ 14y agoNSStrings are opaque--you always call accessor functions and never have access to the low level backing store. The reason they are good is that you can't get data into or out of them without specifying an encoding, which leaves the actual encoding of the backing store as an implementation detail. The fact is, I don't even know (or see documented) that the backing is UTF-16--Apple is free to change that at their whim and no user programs would break.
- brigade 14y agoIt's not documented (presumably) for that very reason. In fact, the opposite is implied by initWithBytesNoCopy:length:encoding:freeWhenDone: - it should be possible right now to have NSStrings with arbitrary internal representations, even if most other creation methods currently convert to UTF16.
- antidoh 14y agoText is maddening, the modern Tower of Babel. Is there a definitive reference, or small handful of references, to learn all that's worth knowing about text, from ASCII to UTF-∞ and beyond?
- njs12345 14y agoJoel Spolsky's 'The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!)' is a good start: http://www.joelonsoftware.com/articles/Unicode.html http://www.joelonsoftware.com/articles/Unicode.html Like a few other specialised fields (cryptography comes to mind) the key takeaway is to use a library and rely on the work of people who know it better than you do and have handled all the subtleties already :)
- haberman 14y agoTotally agree re: UTF-8 vs other Unicode encodings. But are there still still hold-outs who don't like Unicode? Last I heard some CJK users were unhappy about Han Unification: http://en.wikipedia.org/wiki/Han_unification http://en.wikipedia.org/wiki/Han_unification
- lmm 14y agoThe main problem is that it means sort-by-unicode-codepoint puts things in a ridiculous order in japanese/korean. I kind of wish UTF-8 had the latin alphabet in a silly order, so that western programmers would realise they need to use locale-aware sort when sorting strings for display.
- sanxiyn 14y agoThis is false. UTF-8 sorts Korean almost correctly. For practical purposes, you can use sort-by-unicode-codepoint to sort Korean.
- jeffdavis 14y agoI spoke with several Japanese people who said that some valid characters are not representable in Unicode. That means that it's not just a technical problem (expensive sort routines or inefficient encodings) -- it's a semantic problem.
- thristian 14y agoThe way I've heard it explained, there are some historical alternate versions of some characters (A Latin-alphabet equivalent might be the way we sometimes draw "a" with an extra curl across the top, and sometimes without) that have the exact same semantic meaning, and so they were 'unified' to a single code-point. Unfortunately. some people spell their names exclusively with one variant or the other, and Han unification makes that impossible in Unicode.
- adavies42 14y agoisn't the real problem that you can't guarantee correct rendering of ideograph text without specifying fonts? there are japanese kanji that are drawn differently from the chinese hanzi they're descended from, but they're the same from a unicode perspective. imagine if roman, greek, cyrillic, hebrew (aramaic), and ethiopian (ge'ez) were all assigned to the same group of code points and distinguishable only by font--they're all just variants of phoenician, after all....
- erichocean 14y agoThe strangest thing about Unicode (any flavor) is that NULL, aka \0, aka "all zeros" is a valid character. If you claim to support Unicode, you have to support NULL characters; otherwise, you support a subset. I find most OS utilities that "accept" Unicode fail to accept the NULL character. FWIW, UTF-8 has a few invalid characters (characters that can never appear in a valid UTF-8 string). Any one of them could be used as an "end of string" terminator if so desired, for situations where the string length is not known up front. We could even standardize which one (hint hint). I suggest -1 (all 1s). UPDATE: I meant "strange" as in "surprising", especially for those coming from a C background, like me.
- colanderman 14y ago-1 is not a valid Unicode code point. "All 1s" is not adequately defined without saying how many 1s – and Unicode does not specify a maximum bit width. Even if you said "the maximum Unicode code point", that is not all 1s – it is 0x10FFFF.
- erichocean 14y agoThat's the entire point of choosing -1 as an "end of sequence" marker for a UTF-8 string when the length is not known up front. A byte containing all 1s is not valid in any Unicode encoding, so if one appears, you'd know you had hit the end of the string.
- adobriyan 14y agoThis doesn't work well with handling ill-formed sequences. The length is of course known upfront if not of the whole string but at least of the individual small substring.
- colanderman 14y agoOK, I thought you meant a code point containing all 1s. Thanks for clearing that up.
- lubutu 14y ago
- mkup 14y agoI use UTF-8 for transmitted data and disk I/O, and I use UCS-4 (wchar_t on Linux/FreeBSD) for internal representation of strings in my software. I generally agree with this article, but I disagree with it on the point that UTF-8 is the only appropriate encoding for strings stored in memory, and also I disagree on the point wchar_t should be removed from C++ standard or made sizeof 1, as in Android NDK. Let me explain why. In UTF-8 single Unicode character may be encoded in multiple ways. For example NUL (U+0000) can be encoded as 00 or as C0 80. The second encoding is illegal because it's longer than necessary and forbidden by standard, but naive parser may extract NUL out of it. If UTF-8 input was not properly sanitized, or there is a bug in charset converter, this may result in exploit like SQL injection or arbitrary filesystem access or something like that: malicious party can encode not only NUL, but ", /, \ etc this way. Also UTF-8 string can't be cut at arbitrary position. Byte groups (UTF-8 runes) must be processed as a whole, so appear either on left side or on the right side of cut. Reversing of UTF-8 string is tricky, especially when illegal character sequences are present in input string and corresponding code points (U+FFFD) must be preserved in output string. I think UTF-8 for network transmitted data and disk I/O is inevitable, but our software should keep all in-memory strings in UCS-4 only, and take adequate security precautions in all places where conversion between UTF-8 and UCS-4 happens. And sizeof(wchar_t)==4 in GCC ABI is not a design defect, wchar_t exists for a good reason. I admit that sizeof(wchar_t)==2 on Windows is utterly broken.
- ubershmekel 14y agoConcerning "cut at an arbitrary position" actually utf-8 is the only codec that can deterministically continue a broken stream because bytes that start a character are special.
- ruediger 14y ago> Also UTF-8 string can't be cut at arbitrary position. Neither can be any other kind of Unicode string because of Combining Characters. That's why the Unicode standard (or an Annex) recommends algorithms for text segmentation. (And if you really need to cut at a certain length then you can easily backtrack and find the beginning of the sequence by looking for the first byte with the MSB = 0)
- pilif 14y agoReally good article. You'll get nothing from me but heartfelt agreement. I especially liked that the article was giving numbers about how inefficient UTF8 would be to store Asian text (not really apparently). Also insightful, but obvious in hindsight: Not even in utf-32 you can index specific character in constant time due to the various digraphs. The one property I really love about UTF8 is that you get a free consistency check as not every arbitrary byte sequence is a valid UTF8 string. This is a really good help for detecting encoding errors very early (still to this day, applications are known to lie about the encoding of their output). And of course, there's no endianness issue, removing the need for a BOM which makes it possible for tools that operate at byte levels to still do the right job. If only it had better support outside of Unix. For example, try opening a UTF8 encoded CSV file (using characters outside of ASCII of course) in Mac Excel (latest versions. Up until that, it didn't know UTF8 at all) for a WTF experience somewhere between comical and painful. If there is one thing I could criticize about UTF8 then that would be its similarity to ASCII (which is also its greatest strength) causing many applications and APIs to boldly declare UTF8 compatibility when all they really can do is ASCII compatibility and emitting a mess (or blowing up) once they have to deal with code points outside that range. I'm jokingly calling this US-UTF8 when I encounter it (all too often unfortunately), but maybe the proliferation of "cool" characters like what we recently got with Emoji is likely going to help with this over time.
- yuhong 14y agoYea, reminds me of DBCS. UTF-8 however don't use bytes below 0x80 as anything other than an ASCII character, unlike some DBCS encodings such as Shift-JIS.
- _3u10 14y ago"The one property I really love about UTF8 is that you get a free consistency check as not every arbitrary byte sequence is a valid UTF8 string." You don't get this at all using UTF-8. You only get it if you attempt to decode the string which even something like strlen doesn't do. Strlen will happily give you wrong answers about how many characters are in a UTF-8 string all day long and never ever attempt to check the validity of the string. Take your valid UTF-8 and change one of the characters to null, now it doesn't work in many circumstances with 'UTF-8' code. Also, should the free consistency check ever actually work you're in a bigger pickle as you now have to figure out whether the string is wrongly encoded UTF-8 or someone sent you extended ASCII. I did a lot of work with unicode apps. I used to have a series of about 5 strings that I could paste into a 'UNICODE' application and have it invariably break. One was an extended ASCII string that happend to be valid UTF-8 sans BOM :) One was a UTF-8 string with BOM and has 0x00 inside :) (I call this string how to tell if it was written with C) One was a UTF-8 string with a BOM :) One UTF-8 string with a some common latin characters, a couple japanese, and a character outside the BMP. Two UTF-16 strings in LE/BE with and sans BOM.
- scoith 14y agoThat page is misleading when it comes to Japanese text: UTF-8 sucks for Japanese text. UTF-8 and UTF-16 aren't the only two choices within the whole world, which is demonstrated in their choice of encoding Shift-JIS.
- ruediger 14y agoCan you elaborate on that? Why does Unicode suck for Japanese text?
- byuu 14y agoNot only kanji, but also hiragana and katakana (syllabic alphabets) encode to three bytes per character. Shift-JIS can encode all three to two bytes, as well as half-width katakana to one byte per character. However, if size is such a concern (eg for web transmission), text compression neutralizes the perceived benefit of region-specific encodings. Shift-JIS' continued popularity has much more to do with change aversion than it does technical merit.
- khuey 14y agoOn the web ASCII (think HTML tags, CSS stylesheets, etc) typically is a large fraction of CJK pages, so the relative inefficiency of UTF-8 for encoding is less important.
- jeffdavis 14y agoAs I said above, I spoke with several Japanese people who said that some valid characters are not representable in Unicode. Some details can be found here: http://en.wikipedia.org/wiki/Han_unification http://en.wikipedia.org/wiki/Han_unification
- scoith 14y ago@ruediger There's nothing wrong with Unicode. UTF-8 sucks because it ends up taking more space. @byuu No it doesn't. Try compressing a SJIS text using gzip. Then convert it to UTF-8 and do the same thing. With a "perfect" compressor, there shouldn't be any difference since the information contents are the same, but unfortunately we don't have a perfect compression algorithm that hits the theoretical lower bound for compression.
- fleitz 14y agotl:dr; Use UTF-8 when you need to use unicode with legacy APIs, never anywhere else. UNIX isn't UTF-8 because UTF-8 is better, UNIX is UTF-8 because you can pass UTF-8 strings to functions that expect ASCII and it kinda works. This is really the only thing you need to know about UTF-8 and why it's better. There are few pieces of software that don't have to talk to legacy APIs that store strings natively in UTF-8. C# and Java are probably the best examples of software that was engineered from the ground up and thus uses UTF-16 internally because it's much less likely to run into issues like String.length returning 32 yet only containing 31 characters. If you use UTF-8 expect this result anytime a string contains a real genuine apostrophe. "UTF-8 and UTF-32 result the same order when sorted lexicographically. UTF-16 does not." This is complete and utter bullshit, to sort a string lexicographically you need to decode it, if you've decoded the string into UNICODE then they sort the exact same way. There are lots of gotchas for sorting UNICODE strings including normalization because you can write the semantically equivalent strings in unicode multiple ways. eg. ligatures. If you're sorting bit strings that happen to contain UTF-8/32 then you're not sorting lexicographically and your results will be screwed up anyway.
- gwillen 14y ago> decoded the string into UNICODE I think you are quite confused. 1) Unicode is not an acroynm. 2) You cannot "decode into Unicode". I think you mean "decode into codepoints". 3) If that is what you mean, then you are wrong about sorting: Sorting UTF-8 and UTF-32 bytestrings will indeed sort them lexicographically by code point, which was the author's point. No, that will not generally be the sort you _want_; but no amount of 'decoding' will give you the sort you want. For that you need to first normalize, and then follow the collation rules, which don't sort by raw code points at all.
- deleted 14y ago[deleted]
- Tloewald 14y agoYour tl;dr is misleading, doesn't represent the thrust of the article, cherry picks nits, and makes assertions that are contradicted with evidence in the article (e.g. UTF-16 is not fixed length).
- chj 14y agocan not agree more! it will be a much better world if we all use utf8 for external string presentation. i don't care about what your app use internally, but if it generates output, please use utf8.
- cygx 14y agoPersonally, I prefer UTF-8 as well. However, I think this whole debate about choice of encoding gets blown out of proportion. Consider the following diagram: [user-perceived characters] <-+ ^ | | | v | [characters] <-> [grapheme clusters] | ^ ^ | | | | v v | [bytes] <-> [codepoints] [glyphs] <----------+ Choice of encoding only affects the conversion from bytes to codepoints, which is pretty straight-forward: The subtleties lie elsewhere...