5 ms·
Man, UCS-2 is the pits. I still remember fighting with 'slim-builds' of python back in the day. Any critique of unicode while not assuming UTF-8, which allows
by herge 10y ago
Man, UCS-2 is the pits. I still remember fighting with 'slim-builds' of python back in the day.
Any critique of unicode while not assuming UTF-8, which allows for more than 1 million code points) is a bit suspect in my opinion. The biggest point against UTF-8 might be that it takes more space than 'local' encodings for asian languages.
- achamayou 10y agoWhich really isn't a very serious concern at this point, considering how cheap storage and compression are. What's expensive to store are images and sound, from an ever increasing number of devices at an ever higher resolution. The production and storage of text barely registers in comparison.
- imagist 10y agoText size is still relevant due to bandwidth and reliability of transmission. Not everyone has a gigabit internet connection--large portions of the world are still operating on 2G wireless, or even dialup.
- ssalazar 10y agoDefinitely. I would be interested to know of any good statistics on data size of real-world non-European language webpages in UTF-8 vs. UTF-16 and also with/without compression. Much of the markup will be smaller in UTF-8 but actual text content would be smaller in UTF-16.
- creshal 10y agoAnd both end up being fed to optimized-deflate-encoder-of-the-week and have negligible size differences once compressed.
- dom0 10y agoSee comment above parent: > This happens for pure text[nb 2] but actual documents often contain enough spaces and line terminators, numbers (digits 0–9), and HTML or XML or wiki markup characters, that they are shorter in UTF-8. For example, both the Japanese UTF-8 and the Hindi Unicode articles on Wikipedia take more space in UTF-16 than in UTF-8.[nb 3]
- jcranmer 10y agoThere was an experiment (see https://bug416411.bmoattachments.org/attachment.cgi?id=303074 https://bug416411.bmoattachments.org/attachment.cgi?id=30307... for the results). that switched all the UTF-16 internal strings in Firefox to UTF-8. It found that UTF-8 lowered memory usage, even on text-heavy East Asian pages. Keep in mind that things like tag names or attribute names are all interned, so there's no space savings from, say, compressing <img src="">.
- stupidcar 10y agoWrong. Most of the world is on 3G and wifi: https://opensignal.com/reports/2016/08/global-state-of-the-mobile-network/ https://opensignal.com/reports/2016/08/global-state-of-the-m... Many "developing" countries never even deployed 2G and dial-up to any great extent. They were simply too poor to build large-scale telephone network infrastructure. When they did start getting connectivity in the late 1990s and 2000s, they were able to skip straight to the latest generation of technology. This is a pattern we see all over the world — the wealth advantage of "developed" nations is offset by their historical investment in infrastructure that is no longer state-of-the-art. For example, the London Underground has been in operation for over a hundred years. The newest bits are great, but the oldest parts are hamstrung by design decisions made in the Victorian era. Whereas, when China builds a new metro, it's able to build every part of it to modern standards, applying the accumulated knowledge from building those earlier metros.
- imagist 10y agoNo, not wrong. I said: > Not everyone has a gigabit internet connection--large portions of the world are still operating on 2G wireless, or even dialup. Sure, the majority of people have moved over to 3G or better worldwide, but there are still many areas where 2G is more common. We just did a deployment in India[1] which still has more 2G coverage than 3G. Performance on 2G connections was a requirement from our Indian business partners. It's also worth noting that your link contains an implicit bias: it's measuring connections, not people. People with slower connections sometimes simply won't connect at all if your site doesn't perform on their connection, so this is always going to skew toward faster connections. Your link is also pretty vague on the actual statistics--given their claim that the vast majority of countries have > 3G availability 75% of the time, 25% of the majority of countries could not have > 3G availability, and if the vast minority country is India, that's hundreds of millions of people. Yes, the majority of the world is on 3G or better, but the minority can still contain millions and millions of people. [1] http://www.sensorly.com/map/2G-3G/IN/India/Vodafone/gsm_40401#|coverage http://www.sensorly.com/map/2G-3G/IN/India/Vodafone/gsm_4040...
- mjevans 10y agoWikipedia has a summary of comparisons: https://en.wikipedia.org/wiki/UTF-8#Compared_to_UTF-16 https://en.wikipedia.org/wiki/UTF-8#Compared_to_UTF-16 Advantages * Byte encodings and UTF-8 are represented by byte arrays in programs, and often nothing needs to be done to a function when converting from a byte encoding to UTF-8. UTF-16 is represented by 16-bit word arrays, and converting to UTF-16 while maintaining compatibility with existing ASCII-based programs (such as was done with Windows) requires every API and data structure that takes a string to be duplicated, one version accepting byte strings and another version accepting UTF-16. Text encoded in UTF-8 will be smaller than the same text encoded in UTF-16 if there are more code points below U+0080 than in the range U+0800..U+FFFF. This is true for all modern European languages. Most communication and storage was designed for a stream of bytes. A UTF-16 string must use a pair of bytes for each code unit: * * The order of those two bytes becomes an issue and must be specified in the UTF-16 protocol, such as with a byte order mark. * * If an odd number of bytes is missing from UTF-16, the whole rest of the string will be meaningless text. Any bytes missing from UTF-8 will still allow the text to be recovered accurately starting with the next character after the missing bytes. Disadvantages * Characters U+0800 through U+FFFF use three bytes in UTF-8, but only two in UTF-16. As a result, text in (for example) Chinese, Japanese or Hindi will take more space in UTF-8 if there are more of these characters than there are ASCII characters. This happens for pure text[nb 2] but actual documents often contain enough spaces and line terminators, numbers (digits 0–9), and HTML or XML or wiki markup characters, that they are shorter in UTF-8. For example, both the Japanese UTF-8 and the Hindi Unicode articles on Wikipedia take more space in UTF-16 than in UTF-8.[nb 3]
- millstone 10y agoIt doesn't mention the biggest disadvantages of UTF-8 relatives to UTF-16: the existence of non-shortest forms and invalid code units.
- syncsynchalt 10y agoAh, but those are disallowed by spec and in the former case you'd open a security bug against anything that didn't transform it to U+FFFD or similar (eg. it lets you sneak '\0's into C-style strings and '/'s into unix paths). So I'll grant you a point but could match that against similar problems in UTF-16 (bad surrogate pairs, surrogate singletons, BOM bombs, and the same invalid code units).
- PeCaN 10y agoSCSU and BOCU-1 can both help a lot with that last point. They're two standard methods of compressing unicode strings that are primarily in a single language—obviously it's not as good as lz4 or gzip or other general-purpose compression algorithms, but it also doesn't increase the size so is ideal for up to ~a few kb.
- mjevans 10y agoI think I'd prefer transparent stream/file-system compression of text documents for this type of issue.
- rspeer 10y agoSCSU and BOCU-1 were never seriously used, as far as I can tell, except to prove a point back when developers were uncertain about implementing Unicode. Unicode standardizer: "Your software needs to support more than one language at a time. Please use Unicode." Developer: "I don't want to. I've already got a codepage that's designed for the language I care about, and Unicode will make everything take up too much space." Unicode standardizer: "Here's BOCU-1, an encoding of Unicode that compresses monolingual text into nearly the same amount of space as your favorite codepage." Developer: "Uh, thanks, but that's weird." Unicode standardizer: "Yeah, never mind. How about you try this new encoding called UTF-8?" Developer: "Oh, this works really well and I guess it's small enough. I'll use it."
- nabla9 10y agoEncoding text is not the problem with Unicode. The problem is the complexity at the higher abstraction level. At the top level of abstraction is the abstract character aka user perceived character and grapheme clusters (a sequence of coded characters that should be kept together). Mapping from code-points to abstract characters is not total, injective, or surjective. Unicode is not just standard for encoding. It's also standard for representation, and handling text. The exact semantics of UTF-8 string in all cases is something that only few programmers are able to comprehend. I know for sure that I don't and I don't know anyone who does. Interchanging and storing UTF-8 strings and hoping for the best is the standard practice and it works well, but it shows the overreach that Unicode standard is.