5 ms·
If that were the case, wouldn't a developer either be using uchar types or wchar_t which have their own specific variants of character related functions?
by birb007 6y ago
If that were the case, wouldn't a developer either be using uchar types or wchar_t which have their own specific variants of character related functions?
- masklinn 6y agoNo? Extended-ascii does not require wchar (which is, incidentally, utter dreck) and isalnum is defined as taking > an unsigned char or […] EOF so isalnum is "their own specific variants of character related functions" for uchar, and locale-aware.
- birb007 6y agoFair point
- iforgotpassword 6y agoBut what the duck? How does that make sense for Chinese? This is bullshit. Great, now isalnum returns true for ä if I'm using de_DE.latin1 or whatever, but is still nonsense for any non-western language. And the price we pay for that is a messed up pile of garbage that nobody understands and can segfault.
- bonzini 6y agoIt's not just Western languages. It's how text was encoded until the 90s and early 2000s for pretty much every language except Chinese, Japanese and Korean (maybe Thai, I am not sure). That's the "famous" CJK acronym for languages that really did require double byte character sets. For example Russian, Greek, Hebrew, and Indic languages all have alphabets that will fit very easily in 128 codepoints (Vietnamese too, which is Latin but with a large number of diacritics).
- iforgotpassword 6y agoYou're right. I faintly remember these times. Japanese even had an 8 bit charset that simply dropped kanji, iirc. But all the platforms that use glibc have moved to utf8 over a decade ago. So even for western languages that code doesn't make sense anymore.
- bonzini 6y agoNot sure, I still see mail that is not UTF-8 encoded, probably some clients are "optimizing" if they only see an occasional accented letter. And for non-Latin languages UTF-8 is pretty much a two-fold increase in size so there may still be usage of single-byte character sets; legacy DBCS aren't entirely dead in China and Japan either.
- iforgotpassword 6y agoThe default locale for pretty much any language on most Linux distros is utf8 at least. Legacy systems exist, as you point out especially in those CJK countries, I guess the transition to Unicode was much more annoying there. But incidentally those old MBCS charsets are those where isalnum et al never made sense in the first place. I'd argue the usefulness of all that glibc code is very close to zero. Sure, the glibc folks can pat themselves on the back for covering the POSIX spec so well but I really prefer the Linux approach here; follow POSIX where it makes sense but omit the insanity.
- bonzini 6y agoThey would pat themselves on the back for writing useful code when it was useful. They didn't write it in 2010, it dates back to the 80s. The comparison to Linux also makes little sense; talking about locales, Linux still supports Microsoft code pages for the FAT file system.
- iforgotpassword 6y agoThen it is time to consider removing this code. If you need it use an old glibc. Your comparison to Linux makes no sense either. It's not that the code pages themselves are the problem, you can use them for conversion and whatnot. But pretending a collection of functions that is unable to do anything meaningful in the present day and crashes for extra points is fine, because it was written long ago just isn't. The musl solution is correct. If you need to handle strings in a locale-aware manner, use a dedicated lib. Better yet don't use C.
- masklinn 6y ago> But what the duck? How does that make sense for Chinese? It doesn't. > This is bullshit. Welcome to POSIX locales. > Great, now isalnum returns true for ä if I'm using de_DE.latin1 or whatever, but is still nonsense for any non-western language. Technically it's already nonsense for western languages as ISO-8859 is long outdated, and any codepage other than ISO-8859-1 will not fit inside a uchar when decoded from UTF-8. Also there were non-western languages which fit in 8-bit encodings for which isalnum would work fine. > And the price we pay for that is a messed up pile of garbage that nobody understands and can segfault. Welcome to POSIX locales. And also C, where not reading and understanding the implication of every word in the specification means you're bad and therefore deserve everything you get. In this case, POSIX clearly specifies that input values outside of EOF or "unsigned char" is UB.
- iforgotpassword 6y ago:-) I'm a C dev. Systems mostly but I still hate locales.
- viraptor 6y agoAs much as this person? https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f027338b0fab0f5078971fbe https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...