5 ms·
> why UTF-32 didn't catch on > memory wasteful The answer is in the question really. If you've got a big pile of mostly-ascii data, quadrupling memory/storage
by jffry 4y ago
> why UTF-32 didn't catch on
> memory wasteful
The answer is in the question really. If you've got a big pile of mostly-ascii data, quadrupling memory/storage to encode it as UTF32 is going to be a pretty tough sell
- Shish2k 4y agoHow many "big piles of mostly-ascii data" are there though? (Does anyone want to write a script which searches /dev/mem and categorises pages of RAM into ascii-or-not-ascii so we can get some meaningful numbers? :P ) (If you’re doing number-crunching on giant CSVs, maybe I can see it being important, but all the ascii files on my desktop that I can think of are pretty trivial)
- steveklabnik 4y agoPretty much every website, due to HTML being ASCII.
- ravi-delia 4y agoI feel like compression will do a better job than deciding on an encoding scheme ahead of time, no? Once gziped I wouldn't expect a difference
- steveklabnik 4y agoI thought I'd read something about it, but when I googled, what I did find was this old HN comment: https://news.ycombinator.com/item?id=8514519 https://news.ycombinator.com/item?id=8514519 > UTF-8 + gzip is 32% smaller than UTF-32 + gzip using the HN frontpage as corpus.
- JoshTriplett 4y agoYou still have to decompress it on the other end, to actually parse and use it. At which point you have four times the memory usage, unless you turn it into some smaller in-memory encoding...such as UTF-8.
- KMag 4y agoIn theory, yes. In practice, no. At my previous job, I wrote a short Python script that took /usr/dict/words and gzipped it and also converted to UTF-16LE (inserting a null byte before every character) and gzipping that. The information content is the same, but the compressed UTF-16LE ends up a bit bigger. IIRC, the difference was more than 1%, but less than 20%. My use case was to show the flaw in the logic a colleague was using to assert that gzipped JSON should be the same size as gzipped MessagePack for the same data, because the information content was the same. It was a quick 5-minute script without having to deal with coming up with a suitable JSON corpus to convert to MessagePack. Among other things, the zlib compression window only holds half as many characters if your characters are twice as big.
- KMag 4y agoFor anyone still reading, out of curiosity, I reran the experiment on my Debian box: $ </usr/share/dict/words gzip --best | wc -c 261255 $ </usr/share/dict/words iconv -f utf-8 -t utf-16le | gzip --best | wc -c 303404 A bit over a 16% size increase form converting the "wamerican" dictionary to UTF-16LE and then compressing.
- teddyh 4y agoFor the large dictionary it’s a tiny bit worse, but still rounds to 16%: $ < /usr/share/dict/american-english-insane gzip --best | wc --bytes 1778330 $ < /usr/share/dict/american-english-insane iconv -f utf-8 -t utf-16le | gzip --best | wc --bytes 2061457
- vlovich123 4y agoMost encoders let you define filters that restructure the data using out of band knowledge to achieve better rates. For example, if you have an array of floating point numbers, rearranging the exponent and mantissa can yield significant savings if you can arrange for a consecutive run of each separately because the generic compressor doesn’t know anything about the structure of the data. Compressors are great but out of band structural compression/reorganization + compressor will always outperform compressor alone.
- mhink 4y agoI thought HTML5, at least, was UTF-8 by default?
- steveklabnik 4y agoIt is, for the exact reason why I brought it up: if it were in UTF-32, it would take up much more memory. (UTF-8 is compatible with ASCII, but even if it was stored in some other encoding, conceptually it could be in ASCII, you know?)
- layer8 4y agoIf we’d design computing technology from scratch today, we might be using 32-bit bytes, or maybe even 64-bit ones. If memory usage is not a concern, there’s really no need to have smaller units. Our world however runs on 8-bit bytes, so it makes some sense for text to be based on that. But also, consider Base64 in UTF-32-encoded JSON. ;)
- RobotToaster 4y ago>How many "big piles of mostly-ascii data" are there though? You just posted this to one.
- Shish2k 4y agoHacker News is a super-text-heavy site, but even with this extreme example, the favicon alone means that 20% of the page weight is binary data. For any normal site, binary data outweighs ASCII by several orders of magnitude. In either case though, we’re talking about a few kilobytes, which I wouldn’t consider “big” — like even if you wanted to write an HN reader app for your esp32 microcontroller, HN being served in UTF32 instead of UTF8 probably wouldn’t be the biggest obstacle :P
- vlovich123 4y ago> This however is not the case when Japanese text is mixed with ASCII control structures. For instance XML or HTML documents include enough in-line control data that is in the ASCII range that UTF-8 becomes more efficient as a format compared to UTF-16 (before compression). For instance the front page of the Japanese Wikipedia is 92KB in UTF-8 and 166KB in UTF-16. https://lucumr.pocoo.org/2014/1/9/ucs-vs-utf8/ https://lucumr.pocoo.org/2014/1/9/ucs-vs-utf8/ The favicon btw is cached and amortized across all HN pages whereas the text is not. I forget where I read this but about a decade or so I remember reading a paper or watching a video (maybe from the Azul folks?) that looked into JVM memory usage and a good chunk of it was strings and the ucs2 encoding was a problem. That’s why even languages that are nominally utf16/32 as the native type will frequently auto detect and special cases latin1 strings (python, js, etc). The other piece of it is that strings are copied around more and processed differently from images. The knock on effects of utf32 can be quite unfortunate (ie rendering your html document is meaningfully slower which you care about even though by weight your images take longer to transfer and show)
- kristianp 4y ago> That’s why even languages that are nominally utf16/32 as the native type will frequently auto detect and special cases latin1 strings (python, js, etc). I feel that this is worth a blog post in itself. I remember years ago comparing Go and C# when processing some large mostly-ascii files. The C# program was faster, much to my surprise, despite storing strings natively in UTF-16. (I don't remember the implementation details, so it may have been an artifact of my implementations).
- jcranmer 4y ago> How many "big piles of mostly-ascii data" are there though? Well, the use case mentioned in the article is a pretty good one: program source code. Even if you're going to be writing in a foreign language, all of the fancy punctuation and whitespace that does useful stuff in the language ends up being ASCII, and a good hunk of the standard library is likely to have ASCII names for types and functions, etc.
- lmm 4y agoIt doesn't matter, but programming pop culture says it's better to use a "fast" implementation rather than a "slow" one, regardless of what you're doing and what you need.
- mr_toad 4y ago> How many "big piles of mostly-ascii data" Most databases. It might be compressed on disk, built no DBA wants all their column lengths quadrupled.
- deleted 4y ago[deleted]
- mlindner 4y agoJavascript, html, XML, JSON, source code in general of all languages, to name a few. Also it's important to look at the time period. The farther back in time you go the larger a percentage of all data was designed for direct human consumption. (This is why things like binary coded decimal existed over binary.)
- ok123456 4y agoUTF-32 isn't even guaranteed to be a single code point.
- deathanatos 4y agoNo, UTF-32 code units are Unicode scalar values¹, always. They are not grapheme clusters, such as the "family: man, woman, boy" emoji from TFA. ¹which is approximately what I think you're saying here. I.e., you're trying to say that a code point might span multiple UTF-32 code units; that is not correct. (It should be simple to see how a code point, which has the range [0, 0x10FFFF], can always fit into a u32.)
- xdfgh1112 4y ago*grapheme cluster