4 ms·
One advantage of UTF-16 is that unlike UTF-8, very few characters you encounter in Real Life invoke surrogate pairs. So for the cost of a one-bit flag per strin
by Carlfish 17y ago
One advantage of UTF-16 is that unlike UTF-8, very few characters you encounter in Real Life invoke surrogate pairs. So for the cost of a one-bit flag per string you can assume two bytes per character in your string operations for the overwhelmingly general case.
- pmjordan 17y agoLet's face it though, the main string operation in web apps (or most software) is concatenation. I strongly doubt there is any point in converting back and forth between UTF-16 and UTF-8 just for that tiny advantage in addressing individual code points - and even then, not quite 100% of the time. For search-and/or-replace (or anything that can be done with regexes), I'm pretty sure that UTF-16 has no advantage over UTF-8, as you can make the state machine operate on the bytes directly.
- jerf 17y ago"So for the cost of a one-bit flag per string you can assume two bytes per character in your string operations for the overwhelmingly general case." Thank you for that clear and concise explanation of the dangers of using UTF-16. Yes, I know that wasn't your intention, but it was the end result. One of the most dangerous library failures you can have is a function that works 99.99% of the time. Or in this case, 100% of the time on the input the English-speaking developer provides but distinctly less than 100% in the field. In this specific case, you can't actually optimize anything because all your optimizations are bugs. You can't just divide by two for character count; that's not an optimization, it's a bug. You can't just multiply by two for a substring operation, because you might chop a character in half, that's a bug. And so on. You'd need a separate type that indicates you've scanned the string to verify it never has split chars and now you might as well be on UCS-2, and that has its own dangers w.r.t. working 99.99% of the time. Much better to use UTF-8, where the dangers are much more apparent, all you have to do is leave the base ASCII case and you're testing UTF-8. Even I, an English-speaking developer, manage to test that case (once I know it exists, anyhow). There's still ways you can screw up but you're off to a much better start.
- nradov 17y agoRight. From what I have seen, the majority of Java code out there is broken in exactly that way.