5 ms·
No, they don't even agree between engine implementations within the same database server. Generally limits in the database are defined as storage limits for how
by Manfred 4y ago
No, they don't even agree between engine implementations within the same database server. Generally limits in the database are defined as storage limits for however a character is defined. That usually means bytes or codepoints.
- jonhohle 4y agoOr act like MySQL and define a code point as fixed to 3-bytes when choosing UTF-8 .
- bonsaibilly 4y agoThankfully MySQL also offers a non-gimped version of UTF-8 that one should always use in preference to the 3-byte version, but yeah it sucks that it's not the "obvious" version of UTF-8.
- tragomaskhalos 4y agoIs this part of MySQL's policy of "do the thing I've always done, no matter how daft or broken that may be, unless I see an obscure setting telling me to do the new correct thing" ?
- bonsaibilly 4y agoThat'd be my guess, but I don't really know. They just left the "utf8" type as broken 3-byte gibbled UTF-8, and added the "utf8mb4" type and "utf8mb4_unicode_ci" collation for "no, actually, I want UTF-8 for real".
- jonhohle 4y agoIt will be a fun day when Unicode crosses the 5-byte UTF-8 encoding threshold :/
- otabdeveloper4 4y agoIt won't. We settled on using stateful combining characters instead. (Remember when the selling point of switching the world to Unicode was "represent all writing systems with a single stateless 16 bit encoding"? Yeah, well, lol.)
- bonsaibilly 4y agoAnything beyond four bytes is composed of multiple code points, happily
- WJW 4y agoNo the default these days is the saner utf8mb4, if you create a new database on a modern MySQL version. But if you have an old database using the old encoding then upgrading databases doesn't magically update the encoding because some people take backwards compatibility serious.
- bruce343434 4y agoto be fair, there's utfmb4