5 ms·
IMHO that is then a "bug" in the HTML spec which should be "fixed" to speak of extended grapheme clusters instead, which is what users what probably expect.
by Longhanks 4y ago
IMHO that is then a "bug" in the HTML spec which should be "fixed" to speak of extended grapheme clusters instead, which is what users what probably expect.
- WorldMaker 4y agoExcept "maxlength" is often going to be added to fields to deal with storage/database limitations somewhere on the server side and those are almost always going to be in terms of code points, so counting extended grapheme clusters makes it harder for server-side storage maximums to agree with the input. Choosing 16-bit code points as the base keeps consistency with the naive string length counting algorithm of JS, in particular, which has always returned length in 16-bit code points. Sure, it's slightly to the detriment of user experience in terms of how a user expects graphemes to be counted, but it avoids later storage problems.
- IshKebab 4y agoYeah but that assumes your database stores text in UTF-16 which seems fairly unlikely. Consistency with JavaScript does make some sense at least.
- WorldMaker 4y agoUTF-16 seems very likely to me in database fields that cleanly support Unicode. It's the native default (for decades) in SQL Server and Oracle. Up until very recently it was the only Unicode safe character format for MySQL (when a true 8-bit UTF-8 character type was finally added after decades of hacks with broken 7-bit encodings). Postgres is the only SQL database I'm aware of that has had clean UTF-8 support for as long of a time that UTF-16 has been recommended or defaulted in every other SQL database. Obviously there are a lot of databases out there with hacks using things like bad UTF-8 in unsafe column formats not designed for it, of course. But if you are working with a database like that, codepoint counting is possibly among your least worries.
- jfk13 4y agoThe "length" of a string in extended grapheme clusters is not stable across Unicode versions, which seems like a recipe for confusion. The length in code units is unambiguous and constant across versions.