12 ms·
Except almost everyone always means #2. No one asked for strings to be ruined in this way, and this kind of pedantry has caused untold frustration from develope
by zionic 5y ago
Except almost everyone always means #2. No one asked for strings to be ruined in this way, and this kind of pedantry has caused untold frustration from developers who just want their strings to work properly.
If you must expose the underlying byte array do so via a sane access function that returns a typed array.
As for “string length in pixels”, that has absolutely nothing to do with the string itself as that’s determined in the UI layer that ingests the string.
- garmaine 5y ago> Except almost everyone always means #2. I don't think this is true. Certainly there is no string-length library I'm aware of that handles it that way. The usual default these days (correct or not) is #4 -- length is the number of unicode code points.
- wolf550e 5y agoSome languages that had the misfortune to be developed while people thought UCS-2 would be enough define string length as number of UTF-16 code units, which is not even on the list because it's not a useful number.
- garmaine 5y agoYeah, and most languages/libraries which predate Unicode define string length as the number of bytes (#1). That's probably the most common interpretation, actually. But most new implementations count code point, in my experience. I believe this is also the Unicode recommendation when context doesn't determine a different algorithm to read. Been a while since I read those recommendations though, so I could be wrong.
- zinekeller 5y ago> I believe this is also the Unicode recommendation when context doesn't determine a different algorithm to read. Except that emojis are universally two "characters", even those that are encoded as several codepoints. Also, non-composite Korean jamo versus composited jamo.
- garmaine 5y agoLike this: “:)” ? Japanese kana also count as two characters. Which they largely are when romanized, on average. Korean isn’t identical but the information density is approximately the same. Good enough to approximate as such and have a consistent rule.
- chrismorgan 5y agoMy experience is that most new stuff is deliberately choosing and even exposing UTF-8 these days, counting in code units (which is there equivalent to bytes). I’d say that not much counts in code points, and in fact that it’s unequivocally bad to do so, posing a fairly significant compatibility and security hazard (it’s a smuggling vector) if handled even slightly incautiously. Python is the only thing I can think of that actually does count in code points: because its strings are not even potentially ill-formed Unicode strings, but sequences of Unicode code points ('\ud83d\ude41' and '\U001f641' are both allowed, but different—though concerning the security hazard, it protects you in most places, by making encoding explicit and having UTF codecs decline to encode surrogates). String representation, both internal and public, is a thing Python 3 royally blundered on, and although they fixed as much as they could somewhere around 3.4, they can’t fix most of it without breaking compatibility. JavaScript is a fair example of one that mostly counts in UTF-16 code units instead (though strings can be malformed Unicode, containing unmatched pairs). Take a U+1F641 and you get '\ud83d\ude41', and most techniques of looking at the string will look at it by UTF-16 code units—but I’d maintain that it’s incorrect to call it code points, because U+D83D has no distinct identity in the string like it can in Python, and other techniques of looking at the string will prevent you from seeing such surrogates. It would have been better for Python to have real Unicode strings (that is, exclude the surrogate range) and counted in scalar values instead. Better still to have gone all in on UTF-8 rather than optimising for code point access which costs you a lot of performance in the general case while speeding up something that roughly no one should be doing anyway. (I firmly believe that UTF-16 is the worst thing to ever happen to Unicode. Surrogates are a menace that were introduced for what hindsight shows extremely clearly were bad reasons, and which I think should have been fairly obviously bad reasons even when they standardised it, though UTF-8 did come about two years too late when you consider development pipelines.)
- int_19h 5y agoIt's not just about optimizing Unicode access. It's also because system libraries on many common platforms (Windows, macOS) use UTF-16, so if you always store in UTF-8 internally, you have to convert back and forth every time you cross that boundary.
- arcticbull 5y agoThat’s what Rust does too, in the standard library offering #1 and #4, UTF-8 bytes and code points. Because they named a code point type a char, though, folks continue to assume a code point is a grapheme cluster. Which it frequently is, and that makes things worse.
- garmaine 5y agoOnly in some languages.
- tsimionescu 5y agoI do think #4 is worse than #2. There is extremely little useful information one can get by knowing the code points but not the grapheme clusters, and even for text editing these are confusing - if I move the cursor right starting before è, i definitely don't want to end up between e and `.
- masklinn 5y agoIn Swift a character is a grapheme cluster, and so the length (count) of a string is in fact the number of grapheme clusters it contains. The number of codepoints (which is almost always useless but what some languages return) is available through the unicodeScalars view. The number of code units (what most languages actually return, and which can at least be argued to be useful) is available through the corresponding encoded views (utf8 and utf16 properties)
- garmaine 5y agoThe more I learn about swift, the more I like it.
- deleted 5y ago[deleted]
- sillysaurusx 5y agoRiddle me this: name one algorithm that requires #2 string length. I'll name a way that it doesn't matter. The reason this isn't an issue is because it's nearly impossible to think of a legit use case.
- kaetemi 5y agoNot length directly in itself, but the abstraction seems applicable to text editing controls.
- arcticbull 5y agoThat’s both #2 and #3, because Unicode characters can have various widths including zero. Gotta find out how far to move the cursor.
- ygra 5y agoThe Unicode character doesn't really have an intrinsic width; the font will define one (and I think a font could define a glyph even for things like zero-width spaces). But as soon as you want to display text you'll get into all sorts of fun problems and string length is most likely never applicable in any way. At that point the Unicode string becomes a sequence of glyphs from a font with their respective placements and there's no guaranteed relationship between the number of glyphs and the original number of code points, anyway.
- s_gourichon 5y agoSince you ask for an example: moving the cursor in a terminal-based text editor requires #2 string length. That said I understand your point about focusing on what is actually needed.
- yorwba 5y agoMoving the cursor doesn't require string length, it requires finding the new cursor position. You can compute #2 length by moving the cursor one grapheme cluster at a time until you hit the end, but if you only need to move the cursor once, the length is irrelevant. (Also, cursor positions in a terminal-based text editor don't necessarily correspond to graphemes.)
- oefrha 5y agoNo. In fixed width settings, it’s usually very important to distinguish 0-width characters (I’ll use “characters” for grapheme clusters), 1-width characters, and 2-width characters (e.g. CJK characters and emojis). See wcwidth(3).
- DemocracyFTW 5y agoVery much so. And one may add that the implementation of UAX#11 "East Asian Width"[1] was done in a myopic, backwards-oriented way in that it only distinguishes between multiples 0, 1 and 2 of a 'character cell' (conceptually equivalent to the width one needs in monospace typesetting to fit in a Latin letter). There are many Unicode glyphs that would need 3 or more units to fit into a given text. * [1] https://www.unicode.org/reports/tr11/tr11-39.html https://www.unicode.org/reports/tr11/tr11-39.html
- zionic 5y ago0-width characters should get a deprecation warning in Language 2.0, along with silent letters and ones that change pronunciation based on the characters surrounding them.
- dminuoso 5y agoThe primary problem is language/library designers/users believing there must be one true canonical meaning of the word „length“ like you just did, and that „length“ would be the best name for the given interface. In database or more subtly various filesystems code the notion of bytes or codepoints might be more relevant. By the way, what about ASCII control characters? Does carriage return have some intrinsic or clearly well defined notion of „length“ to you? What about digraphs like ij in Dutch? Are they a singular grapheme cluster? Is this locale dependent? Do you have all scripts and cultures in mind?
- frosted-flakes 5y agoA CR is a space-type character. A string containing it has a length of 1.
- dotancohen 5y agoWhitespace is the term. And some clients expect that whitespace is not included in string length. "I asked to put 50 letters in this box, why can I only put 42?" would not be an unexpected complaint when working with clients. Even if you manage to convey that spaces are something funny called "characters", they might not understand that newlines are characters as well. Or emojis.
- jcynix 5y agoCredit card numbers come to mind, printed in letters they are often grouped into four number block separated by whitespace, e.g. "5432 7890 6543 4365" and now try to copy-paste this into a form field of "length" 16. Ok, that's more of a usability issue and many front end developers seem to be rather disconnected from the real world. Phone number entry is an even worse case, but I digress ...
- zinekeller 5y agoThe UK Government (at least those based in GDS) has noted it (https://design-system.service.gov.uk/patterns/payment-card-details/#allow-different-formats https://design-system.service.gov.uk/patterns/payment-card-d...), but some definitely are not good here. Also, hypens (or dashes) aren't popular in the US but (somewhat) popular in the UK!
- rvense 5y ago> No one asked for strings to be ruined in this way Except for the, what, 80% of the world's population who use languages that can't be written with ASCII.
- bregma 5y agoAccording to the CIA, 4.8% of the world's population speaks English as a native language and further references show 13% of the entire global population can speak English at one level or another [0]. For reference, the USA is 4.2% of the global population. The real statement is that no one asked a few self-chosen individual who have never travelled beyond their own valley to ruin text handling in computers like the American Standard Code for Information Interchange has. [0] ttps://www.reference.com/world-view/percentage-world-speaks-english-859e211be5634567
- Armisael16 5y ago> The real statement is that no one asked a few self-chosen individual who have never travelled beyond their own valley to ruin text handling in computers like the American Standard Code for Information Interchange has. That much is objectively false. Language designers made a choice to use it; they could’ve used other systems. Also LBJ mandated that all systems used by the federal government use ASCII starting in 1969. Arguably that tied language designers hands, since easily the largest customer for computers had chosen ASCII.
- ufmace 5y agoIt seems unfortunate now, but I'd argue it was pretty reasonable at the time. Unicode is the best attempt I know of to actually account for all of the various types of weirdness in every language on Earth. It's a huge and complex project, requiring a ton of resources at all levels, and generates a ton of confusion and edge cases, as we can see by various discussions in this whole comments section. Meanwhile, when all of the tech world was being built, computers were so slow and memory-constrained that it was considered quite reasonable to do things like represent years as 2 chars to save a little space. In a world like that, it seems like a bit much to ask anyone to spin up a project to properly account for every language in the world and figure out how to represent them and get the computers of the time to actually do it reasonably fast. ASCII is a pretty ugly hack by any measure, but who could have made something actually better at the time? Not many individual projects could reasonably to more than try to tack on one or two other languages with more ugly hackery. Probably best to go with the English-only hack and kick the can down the road for a real solution. I don't think anyone could have pulled together the resources to build Unicode until computers were effective and ubiquitous enough that people in all nations and cultures were clamoring to use them in their native languages, and computers would have to be fast and powerful enough to actually handle all the data and edge cases as well.
- dotancohen 5y ago> Except almost everyone always means #2. Until the string has to be stored in a database. Or transmitted over HTTP. Or copy-pasted in Windows running Autohotkey. Or stored in a logfile. Or used to authenticate. Or used to authorize. Or used by a human to self-identify. Or encoded. Or encrypted. Or used in an element on a web page. Or sent in an email to 12,000,000 users, some of whom might read it on a Windows 2000 box running Nutscrape. Or sent to a vendor in China. Or sent to a client in Israel. Or sent in an SMS message to 12,000,000 users, some of whom might read it on a Nokia 3310. Or sent to my exwife.
- arcticbull 5y agoOr sorted! That’s its own special hell.
- dotancohen 5y agoOr compared! How did I even forget about that. There is no form of normalization that covers all use cases. Full text search... Oh I want to cry...
- JohnHaugeland 5y agoEasy enough to normalize the terminology for length to mean character count and size to mean byte capacity.
- dotancohen 5y ago"I don't know or care what a character is. You're a character. Now limit this field to exactly 50 letters. Not 51. And not 49." And that client doesn't consider spaces nor punctuation as letters. Numbers count, though. Thankfully at the time emojis were not a consideration.
- arcticbull 5y agoSo you mean grapheme clusters or code points? Do you want to count zero-width characters? More specifically what are you trying to do?
- stephen_g 5y agoHonestly I mostly only want #1 on that list, not #2, since most of my stuff is fairly low level (systems programming, network stuff, etc.). So that’s not the best generalisation.
- Agentlien 5y agoThat's funny. As I read this I kept thinking "#1 is obviously the most relevant except for layout and rendering, where #3 is the important one." But that's because I was thinking about allocation, storage, etc. For logic traversing, comparing, and transforming strings #4 is certainly more useful. Oh, but you said #2? I guess that is important since it's what the user will most easily notice.
- pif 5y ago> No one asked for strings to be ruined in this way They only got ruined for people who could not realize that the word is bigger than their backyard.
- pif 5y ago> Except almost everyone always means #2 Only those who think that the graphical part of an information system represents the whole of software development.