6 ms·
I wonder if there is ever going to be an encoding that replaces UTF-8? Or have we hit on some sort of permanent local maxima (not a global one in the sense that
by benjaminjackman 9y ago
I wonder if there is ever going to be an encoding that replaces UTF-8? Or have we hit on some sort of permanent local maxima (not a global one in the sense that UTF-8 carries the baggage of being backwards compatible with ASCII ... though maybe you could argue that's more of a unicode problem than a UTF8 encoding format one).
At this point UTF-8 seems pretty permanent, what would come along to replace it? And if it is likely to be permanent shouldn't node / javascript in general be moving towards deprecating UCS-2 / UTF-16 and giving first class support to UTF-8?
I saw all this because a couple of years ago I had to write a UTF-8 converter before ScalaJS natively supported for a serialization library I had written. I was kind of surprised that the javascript support was so lacking, luckily writing a UTF-8 encoder/decoder isn't that hard of an endeavor.
- pwdisswordfish 9y agoFunctions like String.prototype.charCodeAt, String.prototype.indexOf, String.prototype.substr and such (which operate on indices into 16-bit-wide strings) will have to be supported somehow, so I doubt it.
- gumby 9y agoOnly matters to the Windows platform, though. EDIT: early morning brain fart on my part, js of course too. Grr.
- pwdisswordfish 9y agoWhat? These are cross-platform features, defined in the ECMAScript spec. Nearly all string-processing code written in JavaScript uses them.
- carussell 9y agoString.fromCodePoint and String#codePointAt were introduced for this reason. Or rather, the time to bring the UTF-8 Everywhere initiative to JS would have been when these two methods were introduced for ES6. Unfortunately, that didn't happen and the committee went with UTF-16 (which is not the same thing as UCS-2—the fact that they differ is why these methods were introduced in the first place).
- pwdisswordfish 9y agocodePointAt still takes the same type of index that charCodeAt does. It doesn't address the issue at all. $ jsc >>> '\u{1f4a9}A'.codePointAt(1).toString(16) dca9 And what was the committee supposed to do with all those string indexing functions that have existed since LiveScript times, remove them? Change their semantics overnight and break countless software in the process? If there's any thing JavaScript has been doing right so far, it's backwards compatibility. For better or for worse.
- carussell 9y agoI thought it would have been clear from my comment that I don't regard the new ES6 methods as flawless.
- twic 9y agoThis doesn't seem insurmountable. Three steps: 1. Introduce new methods which allow access to code points without indexing by UCS-2 code unit. An iterator over code points is the key thing. You could also have some opaque kind of index, with a method to get a list of indices for every code point in the string, and then a method to look up a code point by opaque index. The former could be expensive, as it would involve scanning the string to find where each character starts (although perhaps that could be done lazily?), but the latter should be cheap. The results of the expensive method could be cached. 2. Wait N years for the new methods to be widely adopted. 3. Shift the internal representation of strings to UTF-8. Anyone using the new methods will see similar, or better, performance. Anyone using the old methods will see a drop in performance. Implementations could even dynamically choose between representations based on usage patterns. The fact that each web page is its own little JS universe should make that fairly practical.
- codewiz 9y agoI wonder if there is ever going to be an encoding that replaces UTF-8? Maybe WTF-8? https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/
- tedunangst 9y agoNo.
- tentaTherapist 9y agoSection 1: "WTF-8 is a hack..." I doubt that this is going to be replacing anything except in the narrow use-cases mentioned on the same page.
- fanf2 9y agoI should have published this definition of WTF-8 properly http://people.ds.cam.ac.uk/fanf2/hermes/doc/qsmtp/draft-fanf-wtf8.html http://people.ds.cam.ac.uk/fanf2/hermes/doc/qsmtp/draft-fanf...
- mncharity 9y ago> ever going to be an encoding that replaces UTF-8? Or have we hit on some sort of permanent local maxima Our programming languages, type systems and compilers, are still extremely poor at specifying the properties of types, and at permitting variant implementations which preserve them. We're still struggling with basics, like memory layout for locality (eg, arrays of structures of arrays). And many languages still can't manage multiple dispatch. As we slowly become less crippled, it eventually becomes straightforward to use alternate representations and encodings. For instance, UTF strings with inline descriptive bitmasks is already a thing, as is substring-local encoding. So it seems we needn't be trapped in a local maxima, at least long-term. And that perhaps the future will eventually be more heterogeneous. Perhaps like integers and floats, there's both diversity collapse (big endian dies, ieee 754), and heterogeneity (integer packing, SIMD).
- masklinn 9y ago> I wonder if there is ever going to be an encoding that replaces UTF-8? Possibly if CJK keeps gaining importance there may be a new encoding which provides smaller encoding of these characters rather than the current primacy of western/european character sets.
- userbinator 9y agoThat is already served well by UTF-16. 50% smaller than UTF-8 for CJK, yet still covers the full Unicode range.
- rlanday 9y agoMost commonly-used CJK characters are in the Basic Multilingual Plane, and take 2 bytes to represent in UTF-16 and 3 in UTF-8. So you're only really saving 33%, not 50%. Plus, if you're storing e.g. HTML, all the HTML markup characters go from one byte to two bytes, which is going to offset the CJK advantage to some degree (or maybe even outweigh it). I think most web documents are served compressed anyway, which makes the difference even smaller.
- jack1243star 9y agoNo, Unicode fares terribly for CJK languages, see Han Unification. I can't stress enough how it is a lie that Unicode can represent all languages in use, when it cannot even distinguish Japanese from Chinese.
- jcranmer 9y agoIt can't distinguish between English, French, German, Swedish, Norwegian, Spanish, etc. either. And there's no charset that does that. Do you find that a problem? If so, you're very much in the minority; if not, why distinguish between Japanese and Chinese but not English and French?
- kalleboo 9y agoBecause the characters are different enough that it causes actual day-to-day problems (I see text all the time on printouts/signs where some incorrect font substitution has taken place and you get Chinese instead of Japanese characters). You can (probably correctly) argue that these are really just deficiencies in text editing/markup tools where you can't mark the language of text but the fact that that's required for correct presentation only in CJK languages seems to indicate that it's a problem with this unification in particular. How would you like if there was Latin/Greek/Cyrillic unification? I mean they're all just alphabets that make the same sounds they're all basically the same right?