2 ms·
Yes, but also no. You are correct that there are lots of things that you can't do without grouping codepoints into grapheme clusters. And it is unfortunate th
by LukeShu 3y ago
Yes, but also no.
You are correct that there are lots of things that you can't do without grouping codepoints into grapheme clusters. And it is unfortunate that the Go standard library doesn't have utilities for grouping codepoints into grapheme clusters; the Go team's response to that is largely "don't do those things; even if we had utilities for grouping codepoints into grapheme clusters, those things would be hard to do and you are likely to get it wrong; human-level text is complicated, so treat it as opaque as much as you can."
But also, there are still lots of things you can do with just codepoints. You can tokenize. You can trim whitespace. You can do case-conversions and comparisons. You can transcode to a different encoding (whether that be a different Unicode encoding, or something like a JSON where some codepoints will need to be \uXXXX-encoded).
So yeah, codepoints aren't always what you want, but they're what you want more often than Java's "usually codepoints, but sometimes surrogate pairs". Especially because I suspect that very few Java programmers bother to test their code with surrogate pairs.
- erik_seaberg 3y agoI agree that iterating over UTF-16 code units can never be less painful than iterating over UTF-32 codepoints, but to be fair Java (like Win32) comes from more than fourteen years earlier when we still hoped UCS-2 would be enough for the world.