3 ms·
On a practical matter, it seems like a bad idea to have codepoints that can take up to an arbitrary number of bytes - this just screams buffer overflow problems
by 2shortplanks 11d ago
On a practical matter, it seems like a bad idea to have codepoints that can take up to an arbitrary number of bytes - this just screams buffer overflow problems.
So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.
- flohofwoe 11d agoOTH UTF-8 is just one variable-length stream encoding among many others (RLE, LBE128, etc...).
- DmitryOlshansky 11d agoThe bonus is synchonizing at arbitrary point in stream and that ASCII is UTF-8
- saghm 11d agoUnless I'm misremembering, even UTF-16 is variable. You need to bump up to UTF-32 to get fixed-width.
- cyphar 11d agoEven better, it's arguably both -- surrogate characters are valid codepoint values so technically UTF-16 is fixed-width but programs need to have special handling for surrogate pairs meaning it is practically variable-width. Truly the worst of all worlds.
- mafuy 11d agoCorrect me if I'm wrong, but I think all kinds of UTF, including 16 and 32, support arbitrary length for a single effective character. This would be because you can stack modifications as long as you like.
- ElectricalUnion 11d agoWhat you meant by "single effective character" is grapheme clusters. This whole discussion is about variable sized code points.
- Dylan16807 11d agoThe first comment was kind of iffy when it was also talking about buffers and characters, and focusing on code points is mostly a bad focus. It's worth bringing up so nobody thinks fixed width at a single layer is particularly useful, because other layers will still be variable.
- flohofwoe 10d agoYes, UTF-16 is the worst of all alternatives and should be abolished rather sooner than later. UTF-32 is fixed-width for UNICODE code points, but a single visual character (e.g. a "grapheme cluster") can be built from multiple code points. This is separate from the encoding algorithm though, grapheme clusters are mostly a problem for the high level code working with already decoded text data (text rendering, comparison, sorting etc...).
- Retr0id 11d agoIn regular unicode, a grapheme can be made up of an arbitrary number of codepoints (and thus an arbitrary number of bytes), which does cause issues at times.
- saghm 11d ago> this just screams buffer overflow problems Without endorsing this specific idea, I think maybe after over half a century of C that this argument shouldn't get in the way of a new standard. Pretty much every other language has managed to solve this problem, and the people who write new projects in C/C++ have decided they're not concerned about buffer overflows, so if someone decides to start a new project using something like this (or go out of their way to add support for it to an existing project), that's kind of on them. The rest of computing shouldn't get stuck in 1972 forever.
- drfloyd51 11d agoIt’s not about the language. It’s about the runtime environment. Not everything is fully developed UI running on beefy CPUs with gigs of RAM. Sometimes the environment forces a language choice.
- saghm 11d agoI'm failing to see why an embedded environment would have any need for a new encoding format, which is kind of my point: the types of things that are going to be written in C are not the ones going to be adopting completely new backwards-incompatible standards anyhow. If we refuse to try something based on how it would interact of the ecosystem that would likely never consider adopting something like it in the first place, we're literally fixing our computing to constraints from half a century ago and counting. Nobody who is going to write C would stop because of something like a new Unicode scheme with much larger encoding widths, so why should that be an argument against it happening?
- strenholme 11d ago“completely new backwards-incompatible standards” A reasonable person would assume you’re talking about UTF-8000. It’s not completely new: RFC2279, the original UTF-8 proposal, worked exactly like UTF-8000 for codepoints 31 bits or smaller in size. It’s not backwards-incompatible: UTF-8000 is exactly like UTF-8 for 1, 2, and 3-byte long codepoints, and like UTF-8 codepoints for 4-byte long codepoints with a value of 0x10_ffff or smaller (so all UTF-8 codepoints encoded with the first byte being 0b1111_00xx or starting with the bytes 0b1111_0100 0b1000_xxxx). It’s a backwards compatible way of encoding numbers in UTF-8 larger than 0x10_ffff or (0x7fff_ffff with the original RFC2279 proposal). I agree that C isn’t the best language to start a new programming project in. There are things I don’t like about Rust, mainly that there’s only one implementation of it, but if I were to start a new project needing the speed of a system programming language, it makes a lot of sense.
- Pannoniae 11d agoYou don't have a buffer overflow problem if you read it in a memory-safe way i.e. read it in chunks and realloc when you reach the size of your allocation. What you will have is a potential denial-of-service attack - although this one isn't particularly great because there's zero amplification (they might as well just send garbage into your firewall)
- torgoguys 11d agoDOS in what way? Can you clarify? Thx.
- LorenPechtel 10d agoMemory allocation?? Why? Real world, you'll bump into limits based on fonts long before you'll get a buffer that's too large for the stack and worth using the memory allocator for. I can't see any reason to support more than 2^64 characters and lots of headaches from trying to beyond that. You check your buffer writes and reject the character if it's too long.
- explodes 11d agoLimit the codepoint to the number of atoms in the universe (less than 32 bytes).
- __david__ 11d agoI don’t think there’s harm is speccing out the arbitrary encoding and then having a different spec that references that spec but puts hard limits on it. Many rfcs are like that.