6 ms·
Anyone interested in going whole-hog with this, forking Java, and making char utf8 too? I'm guessing it's not worth it, even aside from legal costs.
by throwawaymobule 3y ago
Anyone interested in going whole-hog with this, forking Java, and making char utf8 too?
I'm guessing it's not worth it, even aside from legal costs.
- tialaramex 3y ago> making char utf8 too? What would that even mean? A forked Java's strings could insist their implementation is UTF-8 encoded bytes, but that's strings, you're talking about char. Do you want char just to be a byte, like in C ? But Java already has a byte type.
- LukeShu 3y agoIn my forked Java, a char would be UTF-32, like in Go. (And String would be UTF-8 encoded bytes, as you say.)
- kiitos 3y agoGo has a concept of a byte, and a concept of a rune, but no concept of a char.
- ElectricalUnion 3y agoA 32 bit "wide char" is both very wasteful under normal use and mostly useless for grapheme cluster (what business people probably actually mean by "character") handling.
- kgeist 3y agoIn Go, strings are UTF8, but when you iterate over UTF8 characters like this: for _, c := range str { } variable "c" is 32-bit. So it's not really wasteful, it fits in a register. It can be wasteful if you declare a slice of characters, like []rune, but I don't remember ever seeing it in practice.
- aardvark179 3y agoIn Java strings are stored in the heap either as compact strings for cases where everything fits in 8 bits, or as UTF-16. Iterating over them using chars will give you those 16 bit things, but there are other methods you can use to iterate over them as code points, just like you can in Go. The old char behaviour may not be what you really want, but it can’t be changed without breaking existing code.
- LukeShu 3y agoRight. But this thread is about "OK, but what if we forked Java and could make breaking changes like that." Plus, I'd bet that most (but of course, not all) code that iterates over a Java String using UTF-16 chars is broken; failing to properly handle surrogate pairs, and should instead use those iterate-over-them-as-codepoints/UTF-32 methods.
- erik_seaberg 3y agoMost of the things you could do with a single codepoint aren’t valid unless you search for a group of unicode.IsMark(c) and process them together with the previous codepoint. Now I’m concerned that I don’t see that in the standard library. https://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries https://unicode.org/reports/tr29/#Grapheme_Cluster_Boundarie...
- LukeShu 3y agoYes, but also no. You are correct that there are lots of things that you can't do without grouping codepoints into grapheme clusters. And it is unfortunate that the Go standard library doesn't have utilities for grouping codepoints into grapheme clusters; the Go team's response to that is largely "don't do those things; even if we had utilities for grouping codepoints into grapheme clusters, those things would be hard to do and you are likely to get it wrong; human-level text is complicated, so treat it as opaque as much as you can." But also, there are still lots of things you can do with just codepoints. You can tokenize. You can trim whitespace. You can do case-conversions and comparisons. You can transcode to a different encoding (whether that be a different Unicode encoding, or something like a JSON where some codepoints will need to be \uXXXX-encoded). So yeah, codepoints aren't always what you want, but they're what you want more often than Java's "usually codepoints, but sometimes surrogate pairs". Especially because I suspect that very few Java programmers bother to test their code with surrogate pairs.
- howinteresting 3y agoMany programming languages have learned the lesson that while a string is logically a list of characters, representing one as such is inefficient.
- kaba0 3y agoGo is not good at learning from lessons.
- morelisp 3y agoI don't understand the point of this comment. What didn't Go learn this time?
- howinteresting 3y agoWhile I won't disagree with you generally, in this case I think Go is doing the right thing.
- hyperpape 3y agoTo elaborate on this point, you could consider reading https://manishearth.github.io/blog/2017/01/14/stop-ascribing-meaning-to-unicode-code-points/ https://manishearth.github.io/blog/2017/01/14/stop-ascribing... and the followup post https://manishearth.github.io/blog/2017/01/15/breaking-our-latin-1-assumptions/ https://manishearth.github.io/blog/2017/01/15/breaking-our-l...
- kaba0 3y agoYou have the exact same problem with UTF-32. It can also only express more complex characters with clusters.
- ElectricalUnion 3y agoIt is not worth it because: * If you're not really serious about using strings, why would you care? * If you're really serious about actually doing things with Strings, you are gonna need to reimplement International Components for Unicode (or something similar in scope), and you have a free, libre and already working one for the "UTF-16 String Java" already. You don't have a working one for your custom fork with custom String handling.
- kevin_thibedeau 3y agoAlready done: J++ begat C#.