4 ms·
Makes me wonder how bad utf-8 in off-latin environments really is: what's the fraction of strings that is "language content", and how much is colons, quotes, de
by usrusr 3y ago
Makes me wonder how bad utf-8 in off-latin environments really is: what's the fraction of strings that is "language content", and how much is colons, quotes, decimal numbers and the like? What fraction of coding cultures embraces home locale script for identifiers, and how many continue on the trajectory set in motion by pre-unicode programming languages? My guess would be that utf-8 is still quite bad, but I wouldn't be entirely surprised if actual numbers turned out considerably less bad than one might expect. And how much of the "utf-8 tax" remains if the actual format of most data at rest and in transit isn't json but gzipped json?
- vlovich123 3y agoExperiments I recall from a while back from browsers and JDK indicated the vast majority of memory is consumed by latin1 strings (even for apps in Asia). Of course, those experiments I recall are something like 20 years old and likely biased (websites were smaller, content was still latin1 dominated). It would be interesting to know the results from Chinese apps and websites today. For websites, the markup used the most space and that stuff remains latin1. For apps I don’t know. But also, text size in memory is typically a joke because the expensive bit is all the multimedia that accompanies text these days. Remember, 100mib is 25million characters even in utf32. That’s 1 hour or so of high quality audio or a few seconds of video. In terms of compression, the smaller initial file will typically compress better so not sure what you mean by “utf-8” tax. UTF8 is not slower to process than utf32 as far as I’m aware except for one small corner case of random codepoint access (but then in UTF text processing, you really shouldn’t be doing that I think but I’m not a utf expert ).