11 ms·
UTF-8 Everywhere (2012)
- voaie 10y agoMay be off-topic, I wonder if anyone is planning a redesign of Unicode for the far future? or is there a better way to handle characters, so we don't require a giant library like ICU?
- jandrese 10y agoWhat would you do differently? Unicode isn't complex because people like things that are hard to understand, it's complex because it took on an exceedingly difficult problem.
- voaie 10y agoGiven more and more custom fonts in the OSes/websites, maybe by using some new APIs, we don't need to specify everything in the Unicode standard. We can design a new font format or just a separate datafile, to store those locale-specific information. The Unicode code points then becomes parking slots for different fonts(with locale info to be registered). And We can use the standard/default datafile to keep the old info about the current unicode standard (say Unicode 8.0). This is just my first thought. Seems that the job of ICU is transfered to the OS or web browser.
- voaie 10y agoI think the Unicode standard should not limit the use of fonts. Instead, let the font or the additonal locale datafile tell us how to deal with those locale issues.
- damienkatz 10y agoICU would still be necessary for collation and case conversion.
- jcranmer 10y agoIf your goal is to eliminate ICU, there's not any change you can realistically make. Unicode has problems, but the most obvious things to fix (CJK unification, precomposed versus combining characters, different semantic characters with completely identical graphs (Angstrom sign versus A-with-circle-above, e.g.)) do not eliminate the need for ICU. Languages are horribly complicated. The Turkish ı/İ issue makes capitalization a locale-dependent thing, and things like German ß/ẞ/ss/SS make case conversion in general mind-boggling. The treatment of diacritics in Latin script for collation purposes differs very heavily between major European languages, so sorting and searching are again locale-dependent. And by the time you're dealing with the locale mess of languages, handling locale-specific number, date, and time representations is pretty much trivial. The need for giant Unicode character tables and CLDR tables, or tables that capture similar information, is quite frankly necessary to handle internationalization to any substantial degree.
- Someone 10y ago"so sorting and searching are again locale-dependent" It's worse. Sorting is dependent on the task at hand. http://userguide.icu-project.org/collation http://userguide.icu-project.org/collation: "For example, in German dictionaries, "öf" would come before "of". In phone books the situation is the exact opposite." That page has lots more 'interesting' cases, for example: "Some French dictionary ordering traditions sort accents in backwards order, from the end of the string. For example, the word "côte" sorts before "coté" because the acute accent on the final "e" is more significant than the circumflex on the "o"." That means that, given two strings s and t such that s sorts before t, you can append characters to t to get u which sorts before s. EDIT (after reading the reply of kelnage): _for some strings s and t_
- kelnage 10y agoNo, I don't think that example does imply that. I interpret it as meaning that for the variants of the same "base word" (i.e. all characters are unaccented) the ordering is defined by the positions of the accents rather than their respective orderings. It says nothing about two words that have different lengths or bases.
- PeterisP 10y agoIf you want to handle characters by anything much simpler than current Unicode, you need to simplify the reality that Unicode describes, changing or eliminating a bunch of major human languages. Not all of them, and not even most of them, but still hundreds of millions of people would need to change how they use their language. It could happen in a century or two, actually, we are seeing some language trends that do favor internationalization and simplification over localization and keeping with linguistic tradition.
- voaie 10y agoRight, there will be less common languages. The faded ones could be kept in the digital world by using special fonts.
- vorg 10y agoSimplication (caused by internationalization) and diversification (caused by localization) are two ends of a spectrum, but languages, both their spoken and written forms, have bounced between those ends throughout history. In a century or two, by the time simplification has succeeded on Earth, the settlers on Titan will rebel with their own graphical symbols for displaying language.
- hackuser 10y ago> we are seeing some language trends that do favor internationalization and simplification over localization and keeping with linguistic tradition. I know you're not necessarily advocating it, but if our cultures change to adapt to our technological limitations, that's the reverse of what I think should be happening - there's a problem with the tech.
- deleted 10y ago[deleted]
- mangix 10y agothis seems specific to Windows. UTF8 is already standard in Linux and the web for example. It's just Microsoft.
- tajen 10y agoI'm on Mac and I've had problems with Chrome sending ajax requests or decoding ajax responses in ISO-8859-1, if I remember well. I had to add "; charset=utf-8" to my headers. I remember it was a browser problem, and I think it was the same for all browsers.
- TazeTSchnitzel 10y agoFor backwards-compatibility's sake, where a web page doesn't specify a character set, browsers will assume the predominant pre-Unicode encoding used in your region.
- scrollaway 10y agoWhy is this still the case? UTF8 is dominant now, wouldn't it make more sense to assume UTF8?
- niftich 10y agoThe older the site, the less likely it is that it will have been updated. Therefore, it's reasonable to assume that newer sites will either declare UTF-8, or can be modified to declare UTF-8, while old sites stay the way they always were, pre-UTF-8. Keeping the backwards-compatibility heuristic the same makes sense.
- TazeTSchnitzel 10y agoOld sites lacked encoding declarations, and old browsers (e.g. early versions of IE) didn't support them. Sites that want UTF-8 can ask for it.
- jfries 10y agoAn interesting suggestion they make is to keep utf-8 also for strings internal to your program. That is, instead of decoding utf-8 on input and encode utf-8 on output, you just keep it encoded the whole time.
- codeulike 10y agoWith or without a BOM?
- mcpherrinm 10y agoPutting a BOM in UTF-8 is just silly. Unlike -16, there's no option for which order you put the bytes in. The only time you'll see a BOM in UTF-8 is in poorly converted UTF-16.
- slavik81 10y agoApparently, Powershell requires a BOM to recognize UTF-8 scripts. https://github.com/chocolatey/choco/wiki/CreatePackages#character-encoding https://github.com/chocolatey/choco/wiki/CreatePackages#char...
- codeulike 10y agoYep, but a lot of MS software will only read UTF-8 correctly if a BOM is present.
- damienkatz 10y agoJoking? BOM is completely unnecessary in UTF8, only useful to losslessly preserve UTF16 text when converting back and forth.
- jcranmer 10y agoWithout. UTF-8 is such a distinctive pattern that if text with high bits set matches UTF-8, it's almost certainly UTF-8. There's no need for a BOM to tell you it's UTF-8 (looking at you, Windows), and it can easily confuse software instead.
- umanwizard 10y agoHuh? What would a BOM in UTF-8 even do? 1-byte objects can't have an internal byte ordering.
- 10y ago
- wcoenen 10y agoIt's interesting how history seems to have repeated itself with UTF-16. With ASCII and its extensions, we had 128 "normal" characters and everything else was exotic text that caused problems. Now with UTF-16, the "normal" characters are the ones in the basic multilingual plane that fit in a single UTF-16 code point.
- mark-r 10y agoIt's worse. With UTF-8, if you're not processing it properly it becomes obvious very quickly with the first accented character you encounter. With UTF-16 you probably won't notice any bugs until someone throws an emoticon at you.
- ridiculous_fish 10y agoUnfortunately not. It's easy to process UTF-8 such that you mishandle certain ill-formed sequences that you are unlikely to encounter accidentally. IIS was hit [1], Apache Tomcat was hit [2], PHP was hit twice [3] [4]. UTF-16 has its own warts, but invalid code units and non-shortest forms are exclusive to UTF-8. [1] http://www.sans.org/security-resources/malwarefaq/wnt-unicode.php http://www.sans.org/security-resources/malwarefaq/wnt-unicod... [2] http://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2008-2938 http://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2008-2938 [3] https://www.cvedetails.com/cve/CVE-2009-5016/ https://www.cvedetails.com/cve/CVE-2009-5016/ [4] https://www.cvedetails.com/cve/CVE-2010-3870/ https://www.cvedetails.com/cve/CVE-2010-3870/
- Animats 10y agoThe Python problem is amusing. Python 3 has three representations of strings internally (1-byte, 2-byte, and 4-byte) and promotes them to a wider form when necessary. This is mostly to support string indexing. It probably would have been better to use UTF-8, and create an index array for the string when necessary. You rarely need to index a string with an integer in Python. FOR loops don't need to. Regular expressions don't need to. Operations that return a position into the string could return an opaque type which acts as a string index. That type should support adding and subtracting integers (at least +1 and -1) by progressing through the string. That would take care of most of the use cases. Attempts to index a string with an int would generate index arrays internally. (Or, for short strings, just start at the beginning every time and count.) Windows and Java have big problems. They really are 16-bit char based. It's not Java's fault; they standardized when Unicode was 16 bits.
- dietrichepp 10y agoThere are a couple cases for string indexing, usually involving parsing or regular expressions. You might want to slice the quotes off of a quoted string, or slice from one match of a regular expression to a match of a different regular expression starting at a different index. These come up infrequently enough that it doesn't make sense to make a better API just for these use cases, but frequently enough that it would be a serious impediment if we didn't do some kind of string indexing. I agree, however, that it's completely irrelevant whether the indexes correspond to code units (i.e. byte offsets in UTF-8) or whether they correspond to code points (how it works in Python currently), as long as we have some way to store, compare, and otherwise manipulate locations within a string. Some Rust developers at one point proposed making string indexes their own (opaque) type, as you suggest, so that they couldn't be confused with integers used for other purposes. The extra complexity of such an API meant that this proposal was never really taken seriously, and it only prevents a small category of programming errors. You might be interested in looking at some string APIs which are mostly without string indexing, like Haskell's Data.Text, which is one of the most well-designed string APIs ever made. https://hackage.haskell.org/package/text-1.2.2.1/docs/Data-Text.html https://hackage.haskell.org/package/text-1.2.2.1/docs/Data-T... As for Windows, my Windows apps use UTF-8 everywhere, and then convert to wchar_t at the last possible moment when interacting with the Windows API. I believe this is what UTF-8 Everywhere suggests.
- yuhong 10y agoI have the feeling that back in 1990, ISO 10646 wanted 32-bit characters but had no software folks on that committee, while the Unicode people was basically software folks but thought that 16-bit was enough (this dates back to the original Unicode proposal from 1988). UTF-8 was only created in 1992, after the software folks rejected the original DIS 10646 in mid-1991.
- IvanK_net 10y agoWhen you create a table in MySQL, a text attribute (VARCHAR etc.) is not encoded in UTF8 by default. I think UTF8 should be the default and only format for storing text attributes in all databases and all other text encodings should be removed from database systems.
- zeta0134 10y agoWe can't even convince Microsoft, Apple, and everything else Unix based to agree on line endings. How on earth are we going to convince everyone that one character encoding format is the only way they should store their data? Annoying as it is to deal with, our history as computer scientists demands that we maintain compatibility with older systems and encoding formats that were once used but are now almost forgotten. If we removed all the other encoding formats (code paths that, while underused, still function perfectly fine) we would lose the ability to parse and manipulate a lot of old data.
- chillacy 10y agoThis article was from 4 years ago. Since then, utf 8 adoption has increased from 68% to 87% of the top 10 million websites on Alexa: https://w3techs.com/technologies/history_overview/character_encoding/ms/y https://w3techs.com/technologies/history_overview/character_...
- Const-me 10y ago_Unicode_ adoption increased to 87%. At the cost of non-Unicode encodings. UTF16 isn’t good enough for web: even for a content in Ukrainian or Hebrew languages, UTF8 saves a sizeable bandwidth because spaces, punctuation marks, newlines, digits, English-inspired HTML tags — in UTF8 they all encode in 1 byte per character, and for the web, bandwidth matters.
- chillacy 10y ago> _Unicode_ adoption increased to 87%. At the cost of non-Unicode encodings. Am I reading that site incorrectly? It says: UTF-8: 87.2%, not unicode. Then down below: " The following character encodings are used by less than 0.1% of the websites" UTF-16 https://w3techs.com/technologies/overview/character_encoding/all https://w3techs.com/technologies/overview/character_encoding...
- Murk 10y agoAfter considering this problem in long detail in the past, I too favoured utf8 at the time. I remember a project (circa 1999) I worked on which was a feature phone HTML 3.4 browser and email client (one of the first). The browser/ip stack handled only ascii/code page characters to begin with. To my surprise it was decided to encode text on the platform using utf-16. Thus the entire code base was converted to use 16 bit code points (UCS-2). On a resource constrained platform (~300k ram IFIRC), better, I think, would have been update the renderer and email client to understand utf8. Nice as it might be to have the idea that utf16, or utf32 were a "character" it is as has been pointed out not the case, and when you look into language you can see how it never can be that simple.
- cm3 10y agoOfftopic, but does anyone know of a way to ensure I don't introduce non-ASCII filenames, to ensure broad portability across systems? I've had to resort to disabling UTF-8 on Linux to achieve that.
- viraptor 10y agoWhat's the use case? Make sure you don't introduce them as a desktop user? As a app developer? (what does the app do?) As a sysadmin with third party unknown apps? You can't really "disable utf-8" on Linux. You can change how things are encoded when displaying or saving. (via locale/lang variables) But if the app wants to create a file named "0xE2 0x98 0x83" (binary version of course), it's still free to do that.
- cm3 10y agoI just don't want garbage file names when sharing a file system between systems that don't agree on the encoding. I was thinking maybe some mount option. I can use ISO-88591 and skip UTF-8. I haven't found a mount option for ext4 or xfs yet.
- viraptor 10y agoThere isn't one. The names in ext4 and xfs are opaque binary with some simple limitations (like null bytes). Encodings simply don't exist at the fs layer. You could probably write some filter using fusefs, but in practice... I think you should configure the servers / clients to agree on encoding instead. Better supported and shouldn't be that much work.
- Koromix 10y agoLinux filesystems are not encoding aware. Paths are just treated as opaque byte strings. However, there is ongoing work add configurable safe filenames to Linux: https://lwn.net/Articles/686789/ https://lwn.net/Articles/686789/ But it won't allow you to force ISO-8859-1 in this form. However you could filter out non-ASCII characters.
- hackuser 10y agoIs there any application where UTF-8 isn't the best choice for long-term (i.e., 20-200 year) forward compatibility?
- niftich 10y agoPlaces and situations where you can't accommodate variable-length encodings. As far as future-proofing, UTF-8 is essentially the new ASCII, in that UTF-8 will remain a backward-compatibility goal for any other format that will succeed it.
- hackuser 10y ago> As far as future-proofing, UTF-8 is essentially the new ASCII, in that UTF-8 will remain a backward-compatibility goal for any other format that will succeed it. Yes, I love that every byte transmitted on the Internet still reserves code points for controlling teletype (or similar) machines.
- wrp 10y agoThis militancy to force everyone to use UTF-8 is bad engineering. I'm thinking of GNOME 3, where you aren't even allowed the option of choosing ASCII as a default setting, only UTF-8 or ISO-8859-x. A default setting is just as important for what it filters out as for what it passes through. I use a lot of older tools on *nix that are ASCII-only, in tool chains that slurp and munge text. If the chain includes any of these UTF-8-only apps, I'm constantly dealing with the problem of invalid ASCII passing through.
- misnome 10y agoI quite like Swift's approach -Characters, where a character can be "An extended grapheme cluster ... a sequence of one or more Unicode scalars that (when combined) produce a single human-readable character.". This seems, in practice to mean things like multibyte entries, modified entries, end up as a single entry. As the trade-off, directly indexing into strings is... Either not possible or discouraged, and often relies on an opaque(?) indexing class. The main weirdness I have encountered so far is that the Regex functions operate only on the old, objective-c method of indexing, so a little swizzling is required to handle things properly.
- Const-me 10y ago> In both UTF-8 and UTF-16 encodings, code points may take up to 4 bytes. Wrong: up to 4 bytes UTF16, and up to 6 bytes UTF8. > Cyrillic, Hebrew and several other popular Unicode blocks are 2 bytes both in UTF-16 and UTF-8. Cyrillic, Hebrew and several other languages still have spaces and punctuation, that take a single byte in UTF8. Now it’s 2016, RAM and storage are cheap and declining, but CPU branch misprediction cost is same 20 cycles and not going to decline. > plain Windows edit control (until Vista) Windows XP is 14 years old, and now in 2016 it’s market share is less then 3%. Who cares what was before Vista? > In C++, there is no way to return Unicode from std::exception::what() other than using UTF-8. The exception that are part of STL don’t return Unicode at all, they are in English. If you throw your custom exceptions, return non-English messages in exception::what() in utf-8, catch std::exception and call what() — you’ll get English error messages for STL-thrown exceptions, and non-English error messages for your custom exceptions. I’m not sure mixing GUI languages in a single app is always a right thing. > First, the application must be compiled as Unicode-aware The oldest visual studio I have installed is 2008 (because I sometimes develop for WinCE). I’ve just created a new C++ console application project, and by default it already Unicode-aware. So, for anyone using Microsoft IDE, this requirement is not a problem.
- d0mine 10y agoModern UTF-8 is limited by 4 bytes (not 6). http://stackoverflow.com/questions/9533258/what-is-the-maximum-number-of-bytes-for-a-utf-8-encoded-character http://stackoverflow.com/questions/9533258/what-is-the-maxim... I haven't checked your other claims but this stands out: > The exception that are part of STL don’t return Unicode at all, they are in English. Do you mean they return the text as bytes using some (likely ASCII) character encoding and all the text characters are in ASCII range? There Ain't No Such Thing As Plain Text. (2003) http://www.joelonsoftware.com/articles/Unicode.html http://www.joelonsoftware.com/articles/Unicode.html
- Const-me 10y ago> Do you mean they return the text as bytes using some (likely ASCII) character encoding and all the text characters are in ASCII range? If you rely on std::exception::what() while building a localizable software, you’ll end with inconsistent GUI language. Because some exceptions (that are part of STL) will return English messages, other exceptions (that aren’t part of STL) will return non-English messages. This means if you’re developing anything localizable, you can’t rely on std::exception::what(). Then why care about it’s prototype?
- douche 10y agoSome days, I imagine a parallel universe, where the ancient Chinese had called ideograms a bad idea, and went on to develop a proper alphabet. Unicode would be pretty much unnecessary.