6 ms·
I am wondering what would happen, if someone swapped out the glibc implementation? Would this uncover the deeper technical reason for the implementation? Would
by buster 6y ago
I am wondering what would happen, if someone swapped out the glibc implementation? Would this uncover the deeper technical reason for the implementation? Would suddenly bugs occur elsewhere?
It probably just shows how old the glibc is and like most software entropy is increasing over time.
- rightbyte 6y agoI have not looked at the implementation but a lookup table would suggest EBCDIC support? Otherwise you would have to do like 7 range checks.
- masklinn 6y agoThe comment block mentions locales so I'd think non-C locales and extended-ascii (though non-ascii is also possible) e.g. in ISO-8859-9, 0xDD is İ (uppercase dotted i) and 0xFD is ı (lowercase dotless i) which you would expect to be an alphabetic letter in the turkic locale.
- bonzini 6y agoOr à á é è í ì ò ó ù ú in pretty much anywhere in Europe other than the England (and also places like Greece, Serbia and Bulgaria which have a full alphabet of their own to handle in isalnum :)). Musl assumes your locale is either C or UTF-8, glibc doesn't. Sure almost everything is UTF-8 these days and has been in the last 15 years, but claiming that the glibc code makes no sense is utterly condescending and I expected ddevault to know better than that. I also wouldn't be surprised if those countries that have a non-Latin alphabet were still using single byte character encodings since UTF-8 would double the size of their text.
- masklinn 6y ago> Or à á é è í ì ò ó ù ú in pretty much anywhere in Europe country other than the England (and also places like Greece, Serbia and Bulgaria which have a full alphabet of their own to handle in isalnum :)). That's more complicated because depending on the language these may or may not be considered separate letter e.g. the french alphabet has 26 letters: while required by the grammar, diacritics and ligatures are "extras". This means isalnum wouldn't necessarily return true for, say, ç in the french locale (I've no idea what it actually does). Turkish however does have 29 letters, and the dotted / dotless i is very much part of the alphabet (so are ç, ş, ğ, ö and ü). > Musl assumes your locale is either C or UTF-8 UTF-8 is an encoding, not a locale. And Musl just plain assumes the C locale, it clearly has no support whatsover for anything else. Which, for what that's worth, is a respectable choice: posix locales are really really bad, so not bothering with them is not necessarily a negative, but at the same time dinging glibc for supporting locales or old platform is not very honest.
- bonzini 6y ago> isalnum wouldn't necessarily return true for, say, ç in the french locale (I've no idea what it actually does). It returns true, otherwise "advance to the next word" or "grep -w" would be utterly broken. ("grep -w ça" would not match anything, and in fact it probably won't match anything under musl). Other European languages also treat characters with diacritics as separate letters, for example the Czech alphabet has 42 letters (one of which is "ch" which doesn't have its own Unicode character as far as I remember; if it did there would be more fun to be had with normalization and title case). > And Musl just plain assumes the C locale, it clearly has no support whatsover for anything else. Yeah what I meant is that if you are reading UTF-8 you won't be passing chars in the 128-255 range to the ctype functions. But musl doesn't implement Unicode wchar_t either.
- masklinn 6y ago> Yeah what I meant is that if you are reading UTF-8 you won't be passing chars in the 128-255 range to the ctype functions. But musl doesn't implement Unicode wchar_t either. The trap is that ctype functions take int, that's exactly the mistake TFAA made: they passed in an integer assuming isalnum would work with codepoints (which doesn't actually make sense since int is only required to be 16 bits).
- mwcampbell 6y ago> But musl doesn't implement Unicode wchar_t either. It looks like it does. iswalpha uses some kind of table to return its response. Is that implementation somehow too simplistic?
- bonzini 6y agoYou're right I looked up only iswalnum and iswdigit, and trusted masklinn on the rest. :-) So I remembered right, musl assumes an encoding of Unicode where 128-255 is invalid.