5 ms·
JS, like Java, now has implementations that will also store strings as Latin1 when the implementation believes it is safe to do so. This results in significant
by STRML 7y ago
JS, like Java, now has implementations that will also store strings as Latin1 when the implementation believes it is safe to do so. This results in significant memory savings [1] in most programs.
1. https://blog.mozilla.org/javascript/2014/07/21/slimmer-and-faster-javascript-strings-in-firefox/ https://blog.mozilla.org/javascript/2014/07/21/slimmer-and-f...
- raverbashing 7y agoI wonder why Latin-1 and not just UTF-8
- happytoexplain 7y agoThey're the same encoding, given only characters covered by Latin-1. So it's just a little more explicit to say Latin-1, as you're specifying not only the encoding, but also the set of characters. Edit: I.e. it would be awkward to say you're implementing "UTF-8, but only for these codepoints". That would be equivalent to implementing Latin-1. Edit: I'm thinking of standard ASCII rather than Latin-1. Above the first 128 code points, UTF-8 switches to two bytes while Latin-1 remains with one byte, so it is much simpler.
- zamadatix 7y agoThis explanation has little to do with why. Latin-1 guarantees each character is coded into a single set of 8 bits, UTF-8 is a variable width encoding. The point of giving an encoding is so it is known how to decode it and a string passed as Latin-1 comes with guarantees about character positions and so on without parsing.
- nmadden 7y agoThat’s not correct. Latin-1 and UTF-8 are both compatible with 7-bit ASCII but they are not the same encoding. For instance é (e acute) is a single byte in latin-1 (0xe9) but is two bytes in UTF-8 (0xc3 0xa9)
- happytoexplain 7y agoYou're right! My bad - the first characters are encoded the same across ASCII, UTF-8, and Latin-1, but the second half of Latin-1 differs from UTF-8. So even just having to support those first 256 code points, we jump into multi-byte UTF-8 territory, meaning complexity over Latin-1.
- SigmundA 7y agoThe article has a whole paragraph dedicated to your question
- happytoexplain 7y agoThe article specifies why they didn't convert everything to UTF-8 - not how they chose the encoding for that subset of lower code points.
- SigmundA 7y agoRead the section it describes exactly why they choose Latin-1, it literally answers the question asked in detail...
- happytoexplain 7y agoAh, you're right, though the first bullet point really threw me off: >Gecko is huge and it uses TwoByte strings in most places. Converting all of Gecko to use UTF8 strings is a much bigger project and has its own risks. But I guess they mean converting it all to support UTF-8 under the described circumstances - not just converting it all to UTF-8.
- Macha 7y agoFrom the article: * Gecko is huge and it uses TwoByte strings in most places. Converting all of Gecko to use UTF8 strings is a much bigger project and has its own risks. As described below, we currently inflate Latin1 strings to TwoByte Gecko strings and that was also a potential performance risk, but inflating Latin1 is much faster than inflating UTF8. * Linear-time indexing: operations like charAt require character indexing to be fast. We discussed solving this by adding a special flag to indicate all characters in the string are ASCII, so that we can still use O(1) indexing in this case. This scheme will only work for ASCII strings, though, so it’s a potential performance risk. An alternative is to have such operations inflate the string from UTF8 to TwoByte, but that’s also not ideal. * Converting SpiderMonkey’s own string algorithms to work on UTF8 would require a lot more work. This includes changing the irregexp regular expression engine we imported from V8 a few months ago (it already had code to handle Latin1 strings).
- ninkendo 7y ago> Linear-time indexing This is great for ASCII when you know there's no such thing as combining characters/etc, but I would like to remind everyone reading this that there's no such thing as "linear indexing" of user-perceived characters in Unicode. User-perceived characters need processing in order to index due to Grapheme clusters potentially using many code points together. (https://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries https://unicode.org/reports/tr29/#Grapheme_Cluster_Boundarie...) For instance, (on my machine, a dark-skinned male teacher) is a combination of these characters: - U+1F468 Man - U+1F3FE Medium-Dark Skin Tone - U+200D Zero Width Joiner - U+1F3EB School And knowing what the byte index of the start/end of that character in a string cannot be done by just multiplying an offset by some multiple of number of bytes per character.
- BeeOnRope 7y agoThe String (and CharSequence) API are already warped around the assumption of underlying UTF-16 storage. In particular `charAt()` refers to a UTF-16 code unit position, and in general all methods [1] that take any kind of index are using an index into an explicit or implicit UTF-16 string. All of these functions need to be fast: certainly at least O(1). So for your narrow encoding you need one that maps 1:1 (in a fixed-length way) to UTF-16. That rules out multi-byte UTF-8. You could of course just restrict it to UTF-8 chars that fit in a single byte – but that's just ASCII! You waste a bit per character. So you might has we'll use a full 8-bit encoding, and Latin-1 is convenient as it covers many more characters likely to be encountered in text that can use single-byte UTF-8 in the first place. It is also convenient in that Latin-1 and Unicode codepoints identical so conversion to and from UTF-16 can be very efficient (in fact there are SIMD instructions which do it exactly). In any case, even in the scenario you could somehow use UTF-8 efficiently, it is likely to be worse in this particular scenario: - They both use 1 byte for code points 0-127 - Latin-1 uses 1 byte for 128-255, while UTF-8 uses 2 - Above 255, UTF-8 uses at least two but the Java hybrid string will be using UTF-16 now, for a common case of 2 bytes. So the only place UTF-8 wins is for the unusual case of characters outside the BMP which are 3 bytes in UTF-8 but 4 bytes in UTF-16. Those aren't common at all, and in many scenarios where you'd have a lot of them UTF-16 would still win because it represents many characters in the BMP in 2 bytes that take 3 in UTF-8. --- [1] One might think there is an exception illustrated by the few methods that mention "code point", like `codePointAt`. Yes, these deal in code points, but their indexes are still indexes into a UTF-16 string. So they are kind of a hybrid API. They let you count N UTF-16 code units (which will represent <= N code points) into a string, and then deal with the code point at that location. That leads to weird cases like asking for the code point starting at the second half of a surrogate pair, which gives you the low surrogate alone.
- jrochkind1 7y agoAh, thanks for explaining this. I was confused and thought that Latin-1 fit in UTF-8 in one byte, but I went and confirmed for myself you are right it does not. Hm. In fact this means though that it isn't just the mistake of UCS-2 that results in people wanting "transparent Latin1 under the hood to save space". Even if we had gone right to UTF-8, there would still roughly the same cost savings there. I can't think of any way to deal with unicode that wouldn't have similar cost savings from "transparently doing non-unicode when possible". Well, I suppose unless you made whatever Java (et al) are doing to mark which strings or portions of strings are actually being stored as Latin1, and made that a Unicode encoding. That would of course have it's own... challenges. If you think "transparently storing as Latin1 when possible" is an ugly hack, you'd probably not be happier with it being made a unicode standard encoding. It turns out being able to represent all written human communication is really hard. Unicode actually does a pretty amazing job of it, balancing lots of trade-offs. (And the fact that it has as wide-spread adoption as it does is in part testament to this; just because you make a standard doesn't mean anyone has to or is going ot use it. Some of the trade-offs unicode balanced were making the adoption curve as easy as possible for existing software. Making ascii valid UTF-8 was, I think, proven by history to be the right decision, although it involved trade-offs... such as Latin1 not fitting in one byte :) ). Overall, Unicode is a really succesful standard, technically and in terms of adoption. But UCS-2 was a mistake, that if it had been avoided would avoid some headache. But, i think, not the headache where Java is motivated to "store some things transparently as Latin1 where possible."
- sehugg 7y agoI vaguely recall making an alternate string class in the 90s that did this, also precomputing hashCode() in the constructor.