6 ms·
Thank God someone has the capacity to attack this issue in depth. UTF8 situation with regards to systems level programming is scary and annoying at same time e
by dingdingdang 3y ago
Thank God someone has the capacity to attack this issue in depth. UTF8 situation with regards to systems level programming is scary and annoying at same time
edit: I note that HN filters out utf8 emojis..
- jagged-chisel 3y agoFalling back to emoticons is always an option ;-)
- deleted 3y ago[deleted]
- vanderZwan 3y agoBack when I was stil single some people unmatched me on dating apps for using them. Great way to filter out people not worth your time, really
- bee_rider 3y agoYou could do that, or you could set your age minimum to 30 or so.
- vanderZwan 3y agoI'm sure it would have been worse with people in their tweens, but it's not exclusive to them weirdly enough. Another good one is being completely honest about your height (mostly just filters out American expats but still). Anyway, the only reason I used emoticons was because I use this keyboard: http://www.exideas.com/ME/index.php http://www.exideas.com/ME/index.php ... but it apparently had some side-effects.
- serentty 3y agoThe inability to represent text in the vast majority of the world’s languages is a far bigger issue than the lack of emoji.
- tialaramex 3y ago> UTF8 situation with regards to systems level programming is scary and annoying at same time These days one of the systems level programming languages (Rust) is natively UTF-8. This would have been a huge source of drama in say, 1993 (when UTF-8 is basically brand new) and maybe even in 2003 (by which point it's clear Unicode is a success, but still conceivable that UTF-16 "wins" in some sense) but in 2023 people just shrug - obviously it's UTF-8, why not?
- layer8 3y agoOn Windows this isn’t obvious at all, as the native encoding is UTF-16, and converting to/from UTF-8 on every Windows API call is not only a pessimization both in runtime and memory usage, but also introduces complications in having to handle possible conversion failures (involving unpaired surrogate characters, in particular). C in principle allows to abstract over the system encoding.
- dundarious 3y agoWhat about compiling programs with `cl -utf-8` and using -A APIs? I know the docs say “If the ANSI code page is configured for UTF-8, -A APIs typically operate in UTF-8”, to which I say, just “typically”?! But in practice is it not sufficient?
- layer8 3y agoWindows UTF-8 support is relatively recent and I have no experience with it, so I don’t know. However I expect Windows to just do the same conversions internally or in the linked runtime, that the program would otherwise have to do by itself. I’d assume that there will be edge cases that such programs then can’t handle, such as UI input and file paths containing unpaired surrogates.
- hermitdev 3y agoThe -A APIs are converted to calls to the -W APIs, IIRC, so you pay that cost everywhere. Not sure if that's universally true anymore, but it has been that way at least on the NT based kernel since...inception? Probably a good question for Raymond Chen. Sadly, I can't seem to find anything relevant in a quick search of the Old New Thing.
- deleted 3y ago[deleted]
- cryptonector 3y agoThere is a simple solution: use UTF-8 in the middle and push all codeset conversions to the edges. And, ideally: deprecate non-UTF-8 locales.