19 ms·
It’s not wrong that "\u{1F926}\u{1F3FC}\u200D\u2642\uFE0F".length == 7 (2019)
- bstsb 1y agoironic that unicode is stripped out the post's title here, making it very much wrong ;) for context, the actual post features an emoji with multiple unicode codepoints in between the quotes
- cmeacham98 1y agoFunny enough I clicked on the post wondering how it could possibly be that a single space was length 7.
- ale42 1y agoMaybe it isn't a space, but a list of invisible Unicode chars...
- yread 1y agoIt could also be a byte length of a 3 byte UTF-8 BOM and then some stupid space character like f09d85b3
- robin_reala 1y agoIt’s U+0020, a standard space character.
- eastbound 1y agoIt can be many Zero-Width Space, or a few Hair-Width Space. You never know, when you don’t know CSS and try to align your pixels with spaces. Some programers should start a trend where 1 tab = 3 hairline-width spaces (smaller than 1 char width). Next up: The <half-br/> tag.
- Moru 1y agoYou laugh but my typewriter could do half-br 40 years ago. Was used for typing super/subscript.
- c12 1y agoI did exactly the same, thinking that maybe it was invisible unicode characters or something I didn't know about.
- timeon 1y agoUnintentional click-bait.
- aaron695 1y ago[dead]
- deleted 1y ago[deleted]
- dang 1y agoOk, we've put Man Facepalming with Light Skin Tone back up there. I failed to find a way to avoid it. Is there a way to represent this string with escaped codepoints? It would be both amusing and in HN's plaintext spirit to do it that way in the title above, but my Unicode is weak.
- Mlller 1y agoThat would be … "\u{1F926}\u{1F3FC}\u200D\u2642\uFE0F".length == 7 … for Javascript.
- NobodyNada 1y agoThat would be "\U0001F926\U0001F3FC\u200D\u2642\uFE0F" in Python's syntax, or "\u{1F926}\u{1F3FC}\u{200D}\u{2642}\u{FE0F}" in Rust or JavaScript. Might be a little long for a title :)
- dang 1y agoThanks! Your second option is almost identical to Mlller's (https://news.ycombinator.com/item?id=44988801 https://news.ycombinator.com/item?id=44988801) but the extra curly braces make it not fit. Seems like they're droppable for characters below U+FFFF, so I've squeezed it in above.
- NobodyNada 1y agoThat works! (The braces are droppable for 16-bit codepoints in JS, but required in Rust.)
- Phelinofist 1y agoBefore it wasn't, about 1h ago it was showing me a proper emoji
- mrheosuper 1y ago>We’ve seen four different lengths so far: Number of UTF-8 code units (17 in this case) Number of UTF-16 code units (7 in this case) Number of UTF-32 code units or Unicode scalar values (5 in this case) Number of extended grapheme clusters (1 in this case) We would not have this problem if we all agree to return number of bytes instead. Edit: My mistake. There would still be inconsistency between different encoding. My point is, if we all decided to report number of bytes that string used instead number of printable characters, we would not have the inconsistency between languages.
- com2kid 1y agoHow would that help? UTF-8, 16, and 32 languages would still report different numbers.
- curtisf 1y ago"number of bytes" is dependent on the text encoding. UTF-8 code units _are_ bytes, which is one of the things that makes UTF-8 very nice and why it has won
- ivanjermakov 1y agoI would say Unicode has won, but not UTF-8. UTF-16 is also widely used due to its efficiency on asian texts.
- minebreaker 1y ago> We would not have this problem if we all agree to return number of bytes instead. I don't understand. It depends on the encoding isn't it?
- charcircuit 1y ago>Number of extended grapheme clusters (1 in this case) Only if you are using a new enough version of unicode. If you were using an older version it is more than 1. As new unicode updates come out, the number of grapheme clusters a string has can change.
- baq 1y ago
- Aissen 1y agoI'd disagree the number of unicode scalars is useless (in the case of python3), but it's a very interesting article nonetheless. Too bad unicode.org decided to break all the URLs in the table at the end.
- darkwater 1y ago(2019) updated in (2022)
- DavidPiper 1y agoI think that string length is one of those things that people (including me) don't realise they never actually want. In a production system, I have never actually wanted string length. I have wanted: - Number of bytes this will be stored as in the DB - Number of monospaced font character blocks this string will take up on the screen - Number of bytes that are actually being stored in memory "String length" is just a proxy for something else, and whenever I'm thinking shallowly enough to want it (small scripts, mostly-ASCII, mostly-English, mostly-obvious failure modes, etc) I like grapheme cluster being the sensible default thing that people probably expect, on average.
- baq 1y agoASCII is very convenient when it fits in the solution space (it’d better be, it was designed for a reason), but in the global international connected computing world it doesn’t fit at all. The problem is all the tutorials, especially low level ones, assume ASCII so 1) you can print something to the console and 2) to avoid mentioning that strings are hard so folks don’t get discouraged. Notably Rust did the correct thing by defining multiple slightly incompatible string types for different purposes in the standard library and regularly gets flak for it.
- eru 1y agoPython 3 deals with this reasonable sensibly, too, I think. They use UTF-8 by default, but allow you to specify other encodings.
- xigoi 1y agoI prefer languages where strings are simply sequences of bytes and you get to decide how to interpret them.
- afiori 1y agoI would like an utf-8 optimized bag of bytes where arbitrary byte operations are possible but the buffer keeps track of whether is it valid utf-8 or not (for every edit of n bytes it should be enough to check about n+8 bytes to validate) then utf-8 then utf-8 encoding/decoding becomes a noop and utf-8 specific apis can check quickly is the string is malformed or not.
- impure 1y agoI learned this recently when I encountered a bug due to cutting an emoji character in two making it unable to render.
- kazinator 1y agoWhy would I want this to be 17, if I'm representing strings as array of code points, rather than UTF-8? TXR Lisp: 1> (len " ") 5 2> (coded-length " ") 17 (Trust me when I say that the emoji was there when I edited the comment.) The second value takes work; we have to go through the code points and add up their UTF-8 lengths. The coded length is not cached.
- troupo 1y agoObligatory, Emoji under the hood https://tonsky.me/blog/emoji/ https://tonsky.me/blog/emoji/
- Sniffnoy 1y agoAnother little thing: The post mentions that tag sequences are only used for the flags of England, Scotland, and Wales. Those are the only ones that are standard (RGI), but because it's clear how the mechanism would work for other subnational entities, some systems support other ones, such as US state flags! I don't recommend using these if you want other people to be able to see them, but...
- spyrja 1y agoI really hate to rant on about this. But the gymnastics required to parse UTF-8 correctly are truly insane. Besides that we now see issues such as invisible glyph injection attacks etc cropping up all over the place due to this crappy so-called "standard". Maybe we should just to go back to the simplicity of ASCII until we can come up with with something better?
- guappa 1y agoSure, I'll just write my own language all weird and look like an illiterate so that you are not inconvenienced.
- eru 1y agoYou could use a standard that always uses eg 4 bytes per character, that is much easier to parse than UTF-8. UTF-8 is so complicated, because it wants to be backwards compatible with ASCII.
- spyrja 1y agoTrue. But then again, backward-compatibility isn't really such a hard to do with ASCII because the MSB is always zero. The problem I think is that the original motivation which ultimately lead to the complications we now see with UTF-8 was based on a desire to save a few bits here and there rather than create a straight-forward standard that was easy to parse. I am actually staring at 60+ lines of fairly pristine code I wrote a few years back that ostensibly passed all tests, only to find out that in fact it does not cover all corner cases. (Could have sworn I read the spec correctly, but apparently not!)
- umajho 1y agoIf you want to get the grapheme length in JavaScript, JavaScript now has Intl.Segmenter[^1][^2]. > [...(new Intl.Segmenter()).segment(THAT_FACEPALM_EMOJI)].length 1 [^1]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Intl/Segmenter/Segmenter https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... [^2]: https://caniuse.com/mdn-javascript_builtins_intl_segmenter_segment https://caniuse.com/mdn-javascript_builtins_intl_segmenter_s...
- tralarpa 1y agoFascinating and annoying problem, indeed. In Java, the correct way to iterate over the characters (Unicode scalar values) of a string is to use the IntStream provided by String::codePoints (since Java 8), but I bet 99.9999% of the existing code uses 16-bit chars.
- zahlman 1y agoThis does not fix the problem. The emoji consists of multiple Unicode characters (in turn represented 1:1 by the integer "code point" values). There is much more to it than the problem of surrogate pairs.
- ivanjermakov 1y agoCodepoint is not cluster and cluster is not character. I bet there is "50 falsehoods about Unicode".
- xg15 1y agoThe article both argues that the "real" length from a user perspective is Extended Grapheme Clusters - and makes a case against using it, because it requires you to store the entire character database and may also change from one Unicode version to the next. Therefore, people should use codepoints for things like length limits or database indexes. But wouldn't this just move the "cause breakage with new Unicode version" problem to a different layer? If a newer Unicode version suddenly defines some sequences to be a single grapheme cluster where there were several ones before and my database index now suddenly points to the middle of that cluster, what would I do? Seems to me, the bigger problem is with backwards compatibility guarantees in Unicode. If the standard is continuously updated and they feel they can just make arbitrary changes to how grapheme clusters work at any time, how is any software that's not "evergreen" (I.e. forces users onto the latest version and pretends older versions don't exist) supposed to deal with that?
- re 1y agoWhat do you mean by "use codepoints for ... database indexes"? I feel like you are drawing conclusions that the essay does not propose or support. (It doesn't say that you should use codepoints for length limits.) > If the standard is continuously updated and they feel they can just make arbitrary changes to how grapheme clusters work at any time, how is any software that's not "evergreen" (I.e. forces users onto the latest version and pretends older versions don't exist) supposed to deal with that? Why would software need to have a permanent, durable mapping between a string and the number of grapheme clusters that it contains?
- xg15 1y agoI was referring to this part, in "Shouldn’t the Nudge Go All the Way to Extended Grapheme Clusters?": "For example, the Unicode version dependency of extended grapheme clusters means that you should never persist indices into a Swift strings and load them back in a future execution of your app, because an intervening Unicode data update may change the meaning of the persisted indices! The Swift string documentation does not warn against this. You might think that this kind of thing is a theoretical issue that will never bite anyone, but even experts in data persistence, the developers of PostgreSQL, managed to make backup restorability dependent on collation order, which may change with glibc updates." You're right it doesn't say "codepoints" as an alternative solution. That was just my assumption as it would be the closest representation that does not depend on the character database. But you could also use code units, bytes, whatever. The problem will be the same if you have to reconstruct the grapheme clusters eventually. > Why would software need to have a permanent, durable mapping between a string and the number of grapheme clusters that it contains? Because splitting a grapheme cluster in half can change its semantics. You don't want that if you e.g. have an index for fulltext search.
- chrismorgan 1y agoPrevious discussions: • https://news.ycombinator.com/item?id=36159443 https://news.ycombinator.com/item?id=36159443 (June 2023, 280 points, 303 comments; title got reemojied!) • https://news.ycombinator.com/item?id=26591373 https://news.ycombinator.com/item?id=26591373 (March 2021, 116 points, 127 comments) • https://news.ycombinator.com/item?id=20914184 https://news.ycombinator.com/item?id=20914184 (September 2019, 230 points, 140 comments) I’m guessing this got posted by one who saw my comment https://news.ycombinator.com/item?id=44976046 https://news.ycombinator.com/item?id=44976046 today, though coincidence is possible. (Previous mention of the URL was 7 months ago.)
- program 1y agoI did post this. I found it by chance, coming from this other post https://tonsky.me/blog/unicode/ https://tonsky.me/blog/unicode/
- Ultimatt 1y agoWorth giving Raku a shout out here... methods do what they say and you write what you mean. Really wish every other language would pinch the Str implementation from here, or at least the design. $ raku Welcome to Rakudo™ v2025.06. Implementing the Raku® Programming Language v6.d. Built on MoarVM version 2025.06. [0] > " ".chars 1 [1] > " ".codes 5 [2] > " ".encode('UTF-8').bytes 17 [3] > " ".NFD.map(*.chr.uniname) (FACE PALM EMOJI MODIFIER FITZPATRICK TYPE-3 ZERO WIDTH JOINER MALE SIGN VARIATION SELECTOR-16)
- pwdisswordfishz 1y agoCall me naive, but I think the length of a space character ought to be one.
- jibal 1y agoRead the article ... the character between the quote marks isn't a space, but HN apparently doesn't support emoji, or at least not that one.
- deleted 1y ago[deleted]
- osener 1y agoPython does an exceptionally bad job. After dragging the community through a 15-year transition to Python 3 in order to "fix" Unicode, we ended up with support that's worse than in languages that simply treat strings as raw bytes. Some other fun examples: https://gist.github.com/ozanmakes/0624e805a13d2cebedfc81ea84aa1edf https://gist.github.com/ozanmakes/0624e805a13d2cebedfc81ea84...
- mid-kid 1y agoYeah I have no idea what is wrong with that. Python simply operates on arrays of codepoints, which are a stable representation that can be converted to a bunch of encodings including "proper" utf-8, as long as all codepoints are representable in that encoding. This also allows you to work with strings that contain arbitrary data falling outside of the unicode spectrum.
- acuozzo 1y ago> Python simply operates on arrays of codepoints But most programmers think in arrays of grapheme clusters, whether they know it or not.
- deathanatos 1y ago> which are a stable representation that can be converted to a bunch of encodings including "proper" utf-8, as long as all codepoints are representable in that encoding. Which, to humor the parent, is also true of raw bytes strings. One of the (valid) points raised by the gist is that `str` is not infallibly encodable to UTF-8, since it can contain values that are not valid Unicode. > This also allows you to work with strings that contain arbitrary data falling outside of the unicode spectrum. If I write, def foo(s: str) -> …: … I want the input string to be Unicode. If I need "Unicode, or maybe with bullshit mixed in", that can be a different type, and then I can take def foo(s: UnicodeWithBullshit) -> …:
- slavik81 1y agoThe Python language developers themselves thought that their code only needed to operate on str and later realized that it needed to handle arbitrary bytes. It's a common mistake. A lot of code was written using str despite users needing it to operate on UnicodeWithBullshit. PEP 383 was a necessary escape hatch to fix countless broken programs.
- jfoster 1y agoI run one of the many online word counting tools (WordCounts.com) which also does character counts. I have noticed that even Google Docs doesn't seem to use grapheme counts and will produce larger than expected counts for strings of emoji. If you want to see a more interesting case than emoji, check out Thai language. In Thai, vowels could appear before, after, above, below, or on many sides of the associated consonants.
- deleted 1y ago[deleted]
- voidmain 1y agoI haven't thought about this deeply, but it seems to me that the evolution of unicode has left it unparseable (into extended grapheme clusters, which I guess are "characters") in a forwards compatible way. If so, it seems like we need a new encoding which actually delimits these (just as utf-8 delimits code points). Then the original sender determines what is a grapheme, and if they don't know, who does?
- dang 1y agoRelated. Others? (Also, anybody know the answer to https://news.ycombinator.com/item?id=44987514 https://news.ycombinator.com/item?id=44987514?) It’s not wrong that " ".length == 7 (2019) - https://news.ycombinator.com/item?id=36159443 https://news.ycombinator.com/item?id=36159443 - June 2023 (303 comments) String length functions for single emoji characters evaluate to greater than 1 - https://news.ycombinator.com/item?id=26591373 https://news.ycombinator.com/item?id=26591373 - March 2021 (127 comments) String Lengths in Unicode - https://news.ycombinator.com/item?id=20914184 https://news.ycombinator.com/item?id=20914184 - Sept 2019 (140 comments)
- andy_xor_andrew 1y agohttps://news.ycombinator.com/item?id=27529697 https://news.ycombinator.com/item?id=27529697
- deleted 1y ago[deleted]
- estimator7292 1y agoStuff like this makes me so glad that in my world strings are ALWAYS ASCII and one char is always one byte. Unicode simply doesn't exist and all string manipulation can be done with a straightforward for loop or whatever. Dealing with wide strings sounds like hell to me. Right up there with timezones. I'm perfectly happy with plain C in the embedded world.
- RcouF1uZ4gsC 1y agoThat English can be well represented with ASCII may have contributed to America becoming an early computing powerhouse. You could actually do things like processing and sorting and doing case insensitive comparisons on data likes names and addresses very cheaply.
- zahlman 1y agoThere's an awful lot of text in here but I'm not seeing a coherent argument that Python's approach is the worst, despite the author's assertion. It especially makes no sense to me that counting the characters the implementation actually uses should be worse than counting UTF-16 code units, for an implementation that doesn't use surrogate pairs (and in fact only uses those code units to store out-of-band data via the "surrogateescape" error handler, or explicitly requested characters. N.B.: Lone surrogates are still valid characters, even though a sequence containing them is not a valid string.) JavaScript is compelled to count UTF-16 code units because it actually does use UTF-16. Python's flexible string representation is a space optimization; it still fundamentally represents strings as a sequence of characters, without using the surrogate-pair system.
- deathanatos 1y ago> JavaScript is compelled to count UTF-16 code units because it actually does use UTF-16. Python's flexible string representation is a space optimization; it still fundamentally represents strings as a sequence of characters, without using the surrogate-pair system. Python's flexible string system has nothing to do with this. Python could easily have had len() return the byte count, even the USV count, or other vastly more meaningful metrics than "5", whose unit is so disastrous I can't put a name to it. It's not bytes, it's not UTF-16 code units, it's not anything meaningful, and that's the problem. In particular, the USV count would have been made easy (O(1) easy!) by Python's flexible string representation. You're handwaving it away in your writing by calling it a "character in the implementation", but what is a character? It's not a character in any sense a normal human would recognize — like a grapheme cluster — as I think if I asked a human "how many characters is <imagine this is man with skin tone face palming>?", they'd probably say "well, … IDK if it's really a character, but 1, I suppose?" …but "5" or "7"? Where do those even come from? An astute person might like "Oh, perhaps that takes more than one byte, is that it's size in memory?" Nope. Again: "character in the implementation" is a meaningless concept. We've assigned words to a thing to make it sound meaningful, but that is like definitionally begging the question here.
- zahlman 1y ago
- pron 1y agoIn Java, " ".codePoints().count() ==> 5 " ".chars().count() ==> 7 " ".getBytes(UTF_8).length ==> 17 (HN doesn't render the emoji in comments, it seems)
- TacticalCoder 1y ago[dead]
- deleted 1y ago[deleted]
- shirro 1y agoGrapheme clustering does my head in. I just want to delete the character to the left of the cursor damn it.
- koliber 1y agoI love how the title of this submission is changing every time I come back to HN. At first there was an empty space between the double quotes. This made me click and read the article because it was surprising that the length of a space would be 7. Then the actual emoji appeared and the title finally made sense. Now I see escaped \u{…} characters spelled out and it’s just ridiculous. Can’t wait to come back tomorrow to see what it will be then.
- lovich 1y agoThis article could have well have been named "Falsehoods programmers believe about strings"
- TeMPOraL 1y agoOr, to address GP's concerns more directly, "Falsehoods programmers believe about Unicode filtering in Hacker News submission titles and comments".
- rendx 1y ago(Original renaming thread: https://news.ycombinator.com/item?id=44981808 https://news.ycombinator.com/item?id=44981808)
- pseufaux 1y agoCan anyone recommend a good intro to understanding string encoding article?
- torstenvl 1y agohttps://www.joelonsoftware.com/2003/10/08/the-absolute-minimum-every-software-developer-absolutely-positively-must-know-about-unicode-and-character-sets-no-excuses/ https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
- pseufaux 1y agoHaha, this is fantastic. > So I have an announcement to make: if you are a programmer working in 2003 and you don’t know the basics of characters, character sets, encodings, and Unicode, and I catch you, I’m going to punish you by making you peel onions for 6 months in a submarine. I swear I will. Thank you!
- pseufaux 1y agoAfter reading this, I did a search for mentions of this article and found this StackOverflow gem. The top answer basically picks up where the JoelOnSoftware article leaves off and filled in the rest of the blanks for me. https://stackoverflow.com/questions/2241348/what-are-unicode-utf-8-and-utf-16 https://stackoverflow.com/questions/2241348/what-are-unicode... Still have more reading to do and a lot to learn but this was super informative, so thank you internet stranger.
- Mlller 1y agoThe article nearly equivocates “Rather Useless” and “unambiguously the worst”. Python3 seems more coherent to me than the article's argument: 1. Python3 plainly distinguishes between a string and a sequence of bytes. The function `len`, as a built-in, gives the most straightforward count: for any set or sequence of items, it counts the number of these items. 2. For a sequence of bytes, it counts the number of bytes. Taking this face-palming half-pale male hodgepodge and encoding it according to UTF-8, we get 17 bytes. Thus `len("\U0001F926\U0001F3FC\u200D\u2642\uFE0F".encode(encoding = "utf-8")) == 17`. 3. After bytes, the most basic entities are Unicode code points. A Python3 string is a sequence of Unicode code points. So for a Python3 string, `len` should give the number of Unicode code points. Thus `len("\U0001F926\U0001F3FC\u200D\u2642\uFE0F") == 5`. Anything more is and should be beyond the purview of the simple built-in `len`: 4. Grapheme clusters are complicated and nearly as arbitrary as code points, hence there are “legacy grapheme clusters” – the grapheme clusters of older Unicode versions, because they changed – and “tailored grapheme clusters”, which may be needed “for specific locales and other customizations”, and of course the default “extended grapheme clusters”, which are only “a best-effort approximation” to “what a typical user might think of as a “character”.” Cf. https://www.unicode.org/reports/tr29 https://www.unicode.org/reports/tr29 Of course, there are very few use cases for knowing the number of code points, but are there really much more for the number (NB: the number) of grapheme clusters? Anyway, the great module https://pypi.org/project/regex/ https://pypi.org/project/regex/ supports “Matching a single grapheme \X”. So: len(regex.findall(r"\X", "\U0001F926\U0001F3FC\u200D\u2642\uFE0F")) == 1 5. The space a sequence of code points will occupy on the screen: certainly useful but at least dependent on the typeface that will be used for rendering and hence certainly beyond the purview of a simple function.