3 ms·
Why is it the worst way though?
by tehsauce 3y ago
Why is it the worst way though?
- deleted 3y ago[deleted]
- imron 3y ago> Note about Python 3 added on 2019-09-09: Originally this article claimed that Python 3 guaranteed UTF-32 validity. This was in error. Python 3 guarantees that the units of the string stay within the Unicode code point range but does not guarantee the absence of surrogates. It not only allows unpaired surrogates, which might be explained by wishing to be compatible with the value space of potentially-invalid UTF-16, but Python 3 allows materializing even surrogate pairs, which is a truly bizarre design. The previous conclusions stand with the added conclusion that Python 3 is even more messed up than I thought!
- markmark 3y agoPerhaps the part about it needing lookups of the unicode database and being dependent on the version of the database used?
- lmm 3y agoThat's not true though. It just counts the number of code units, that's not version dependent. It's certainly no worse than counting the number of UTF-16 points (I'd argue it's better since it's less arbitrary - whether something is a unicode scalar is a design decision, whether something is in the BMP or not is mostly an accident of implementation).
- masklinn 3y agoBecause it’s never actually useful. You can’t use that information to know how much actual space it takes (in storage) as nobody sane stores UTF-32, you can’t use it to know much much logical space it takes (aka the user’s interpretation), you can’t use it to know how much visual space it takes (not that you can ever get that), and you can’t use it to segment or process the text. A length in codepoints gives you nothing that’s really actionable, at least not that you’d need outside of a context where you could easily obtain it otherwise.
- L3viathan 3y agoIt is useful: When iterating over a string in Python (which I hope you agree _is_ useful?), you get that many parts.
- chrismorgan 3y ago> When iterating over a string in Python (which I hope you agree _is_ useful?) Not often. There’s almost nothing useful you can correctly do with a sequence of code points.
- masklinn 3y ago> It is useful: When iterating over a string in Python (which I hope you agree _is_ useful?), you get that many parts. That’s… not useful? I can’t say I remember ever caring knowing how many items I would be getting during an iteration[0]. If I want to set an iteration limit I can just… do that, using `islice` or some such. [0] in python anyway, in lower level language there can be a utility in order to pre-allocate an output collection
- deleted 3y ago[deleted]
- boxed 3y agoBut 100% of all those complaints also apply to JS, except that UTF-16 is just much more stupid than UTF-32?
- globular-toast 3y agoWhat is useful? What should it be? Don't say bytes because there is already an idiomatic way to get bytes: `len(bytes(s, enc))` which is both more correct and explicit. Maybe it should just return None because the only useful thing is probably how much "space" it occupies on screen in a fixed-width font, but that's too difficult to know.