7 ms·
Maybe not wrong, but it's the worst option. 5 is the number of code points, and 17 is the number of bytes. Both are reasonable answers. 7 is the number of cod
by dahfizz 3y ago
Maybe not wrong, but it's the worst option.
5 is the number of code points, and 17 is the number of bytes. Both are reasonable answers.
7 is the number of code units for utf-16. Seems like the least useful option.
- Dylan16807 3y agoThat means 7 is also a measure of bytes, just slightly more awkward. So it's roughly on par with 17. For 5, the idea is that while you might want to iterate code points, the total number of code points is less useful than either grapheme count or byte count. I think that argument makes sense.
- lmm 3y ago> That means 7 is also a measure of bytes, just slightly more awkward. It's not a real measure of bytes though. It's the count of bytes in an encoding scheme that is (probably) neither what you use to communicate with the outside world nor what your language runtime uses. (And certainly it's no better than 5, since that's also a measure of bytes in a particular encoding).
- MobiusHorizons 3y agoThe JavaScript language forces utf16 (whether or not v8 uses that representation under the hood). For instance if you want to substring the indexes you pass are for utf16 codepoints
- lmm 3y agoSure, but arguing that that's a good reason for length to count utf16 is purely circular.
- Dylan16807 3y agoLots of systems use UTF-16 internally and externally. Counting bytes in UTF-16 is, on average, almost as useful as counting bytes in UTF-8. I don't think just about anything communicates in UTF-32. 5 is basically just a codepoint count, and as such I don't think its usefulness rating should be between the byte counts.
- mjevans 3y agoOnly Windows and Java come to mind - and BOTH of those are insane for sticking to it when the entire rest of the world has moved on.
- masklinn 3y agoWindows, Java, C#, javascript, a surprising number of XML documents (though less so as time marches on thankfully), ICU I think uses UTF-16 internally (for the same historical reasons as the other 4), JOLIET file names are UCS2, some phones interpret “16-bit” SMS as UTF-16 (the spec says UCS2). > and BOTH of those are insane for sticking to it They don’t really have much of a choice because they exposed those semantics as part of the string interface (or for Windows the interaction is slow low level it can’t be hidden), they have performance guarantees and behaviours which matches that. It’s also why Python uses UTF-32, and went through the entire PEP-393 / FS complication to try and stop blowing up memory left and right: the core team considered that switching strings to UTF8 was a bridge too far. There are approximate solutions, but they come with their own costs and complications (e.g. pypy uses UTF8 strings with lazily constructed indices to emulate UTF-32 strings).
- mjevans 3y agoI'm not a Windows based programmer, but couldn't they leave the old API's in place, but make UTF-8 safe versions available for everyone and switch to that... E.G. with Win 11?
- Joker_vD 3y agoYou can set the system codepage to CP_UTF8 since Win 10, I guess, although IIRC it still doesn't work for input. But a) there is a lot of programs using A() functions that don't expect that and break in subtle ways, e.g. DBCS-encoding-aware programs suddenly break because they don't expect a codepoint to span for more than 2 bytes; b) most of the sanely written programs either use UTF-16 explicitly, or use UTF-8 internally and convert between UTF-8 and UTF-16 before/after calling W() functions.
- temac 3y agoI think that argument makes as much sense as saying that an engine is less useful than a car. And pretending that engine.weight should return the weight of the car.
- Joker_vD 3y ago7 cannot be a measure of bytes because a UTF-16 point takes 2 bytes, so the number has to be even. Did you mean 14?
- Dylan16807 3y agoI meant to write 7. That's why I said "measure of" instead of "number of", because you need to multiply by 2.
- MobiusHorizons 3y agoIt makes just as much sense as 17 (for utf8) in a JavaScript context, where charCodeAt(i) returns a utf-16 code point, and strings at least behave as though the implementation uses an array of uint16_t for the storage. Utf 16 is definitely not my favorite representation, but given that context (which the language imposes) 7 is an important number to be able to know.
- chrismorgan 3y agoCorrection: “utf-16 code point” should read “UTF-16 code unit”.
- deleted 3y ago[deleted]
- globular-toast 3y agoAlso it's trivial to get the number of bytes in Python if that's what is wanted: len(bytes(" ", "utf8")) == 17 This is really the only sane way and makes it explicit which encoding you are using.