5 ms·
UTF-16 is one of these ill-fated developments that curse some languages & platforms (WinNT incl Win10, WinAPI32, Java, Flash, JS, Python 3) to his day. compa
by frik 9y ago
UTF-16 is one of these ill-fated developments that curse some languages & platforms (WinNT incl Win10, WinAPI32, Java, Flash, JS, Python 3) to his day.
compareTo uses 0x19, which means doing the “equal each”
(aka string comparison) operation across 8 unsigned words
(thanks UTF-16!) with a negated result. This monster of an
instruction takes in 4 registers of input:
- RyanZAG 9y agoJava uses UTF8 in latest release for strings without special characters. http://www.baeldung.com/java-9-compact-string http://www.baeldung.com/java-9-compact-string UTF16 is not really a curse for languages that require it. String operations in non-English languages are very fast because of it, and most software these days has to deal with localization.
- lmm 9y agoUTF-16 is the worst of all worlds: it's less efficient than UTF8 for most use cases, requires you to think about endianness, but is still a variable-length encoding. (And the cases that require variable-length encoding are rarer than they are for UTF-8, meaning you're less likely to hit them in testing)
- oblio 9y agoWell, not necessarily a curse, but a suboptimal solution. Are there any situations where UTF16 is a clear upgrade over UTF8?
- Const-me 9y agoAny non-Latin string operations. While technically UTF16 is variable length, 99.99% cases use single word per character. I.e. on modern hardware with branch prediction and speculative execution, these branches don't affect speed. With UTF8, CPU mispredicts branches all the time because spaces, punctuations and newlines are single bytes even in non Latin-1 text.
- vardump 9y agoI think the most common operations are comparisons for equality and copying anyways. UTF-8 is faster for those. I tried out how fast I could make UTF-8 strlen, with an assumption of a valid UTF-8 string. The routine ran at 18 GB/s on a single core using SSE. > With UTF8, CPU mispredicts branches all the time because spaces, punctuations and newlines are single bytes I don't understand this sentence. Why would there be any more mispredictions because of those being single bytes? These days code is so often bandwidth limited if anything, so smaller data helps.
- HelloNurse 9y agoIn non-Latin text, if most characters are 2 bytes but a large minority are 1 byte, the branch prediction in charge of guessing between the different codepoint representation lengths expects 2 bytes and fails very often. Speculative execution (counting in two or three ways simultaneously) might mitigate the performance hit.
- vardump 9y ago> In non-Latin text, if most characters are 2 bytes but a large minority are 1 byte, the branch prediction in charge of guessing between the different codepoint representation lengths expects 2 bytes and fails very often You wouldn't want to process a single code point (or unit) at a time anyways, but 16, 32 or 64 code units (or bytes) at once. That UTF-8 strlen I wrote had no mispredicts, because it was vectored. Indexing is slow, but the difference to UTF-16 is not significant. I guess locale based comparisons or case insensitive operations could be slow, but then again, they'll need a slow array lookup anyways. Which string operation(s) are you talking about?
- legulere 9y agoYou don't need to check the representation doing anything specifically with spaces or newlines. All 0x0A bytes are newline characters in UTF8 and all 0x20 bytes are spaces. The only place you really need to decode UTF8 characters is when you convert it to another format (which you hopefully won't need to do anymore in the far future) or display it (where the decoding is a minuscule factor in performance)
- nwellnhof 9y agoNo, it uses the fixed-width LATIN1 (ISO-8859-1) encoding for compact strings. It wouldn't make much sense to use another variable-width encoding like UTF-8.
- vardump 9y agoThat's path dependence [0]. When all of those were conceived in the nineties, 2-byte UCS-2 seemed to be enough to store all unicode code points. UTF-16 came only later, once it was clear 65535 code points is too few. Had those languages been designed in last 10 years, all of them would pick UTF-8 as their code point format. [0]: https://en.wikipedia.org/wiki/Path_dependence https://en.wikipedia.org/wiki/Path_dependence
- kevingadd 9y agoSome JavaScript runtimes (Firefox's Spidermonkey for one) have an optimization that stores some strings in single-byte format where possible to mitigate the cost of the awful original choice to use UCS-2 for JS strings. I expect some other runtimes do this too, but I don't know any off-hand. IIRC this was motivated by Firefox OS (strings eat up a lot of RAM on memory-starved $50 smartphones) but it pays off on desktops too.
- pcwalton 9y agoV8 and JavaScriptCore do it, I believe.
- BeeOnRope 9y agoJava has taken a couple different shots at this, going back a decade or more, and the newer option is currently enabled in Java 9. Some background: https://stackoverflow.com/q/8833385/149138 https://stackoverflow.com/q/8833385/149138
- ubernostrum 9y agoPython as of 3.3 uses any of three different internal storage mechanisms for strings: 1-byte (latin-1), 2-byte (UCS-2) or 4-byte (UCS-4) depending on the width of the highest code point in the string. This allows the internal storage to always be fixed-width, while still saving space for strings which contain, say, only code points representable in a single byte. Prior to 3.3, the internal storage of Unicode was determined by a flag during compilation of the interpreter; a "narrow" compiled interpreter would use 2-byte strings with surrogate pairs for non-BMP code points, and a "wide" compiled interpreter would use 4-byte strings.
- maxerickson 9y agoWhat about https://www.python.org/dev/peps/pep-0393/ https://www.python.org/dev/peps/pep-0393/ ?
- banthar 9y agoPEP-393 is a stupid compromise. They couldn't choose between UCS-2 and UCS-4, so they are using both. They are wasting tons of CPU cycles converting between them and single character outside of range doubles the size of string. I don't fully understand the use case for extracting codepoints from strings, but they could have just added Java-like: codePoints and keep returning code units from old methods. This is CPU and memory efficient and 100% backwards compatible. I think the problem is the same could have been done in Python 2 (with UTF-8) that would mean less reasons for Python 3.