4 ms·
It's somewhat common to see videogames issue a patch shortly after release where they fix crashes due to non-ASCII Windows usernames or non-English locales. I'm
by supernes 5y ago
It's somewhat common to see videogames issue a patch shortly after release where they fix crashes due to non-ASCII Windows usernames or non-English locales. I'm not sure what the root cause of the confusion is, other than text strings being hard in general.
- jerf 5y agoIt's easy to think the answer is "just UTF-8 everything" but unfortunately the long and twisty history of filesystems means that's not the correct answer, and the "correct answer" is really hard to write down quickly. If you never display the filename, the answer is to treat existing filenames as bags of bytes, but that breaks down as soon as you need to display them, or if you need to manipulate them by appending unicode to them, in which case you have to decide on an encoding. Unicode encodings tend to mangle non-Unicode values because they're specified to replace whatever they can't understand with a particular Unicode character, usually represented as a diamond with an inverted ? inside of it. There's some obscure solutions to this problem, like https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/ (which includes discussion of the 16 bit encodings you need for Windows), but I haven't seen broad support for them. You need a deliberately "noncompliant" encoding/decoding system that doesn't replace unknown characters with replacement characters. Fortunately, compliant systems are becoming more and more popular and available. Unfortunately, that can make file name handling harder than when you had a non-Unicode-compliant handling system for your strings.
- account42 5y ago> If you never display the filename, the answer is to treat existing filenames as bags of bytes, but that breaks down as soon as you need to display them, or if you need to manipulate them by appending unicode to them, in which case you have to decide on an encoding. No you don't. On Windows you treat paths as a u16'\' an/or u16'/'-separated sequences of uint16_t. On Unix it's a '/'-separated sequence of bytes. If you want to display, you need to decode, but for display only - so errors should use replacement characters as a graceful failure. For appending you encode your string and then append the bytes. Never do you decode externally provided paths for the purpose of manipulation. > There's some obscure solutions to this problem, like https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/ (which includes discussion of the 16 bit encodings you need for Windows) It's relatively new, but has wide enough adoption cosidering - e.g. it's what Rust uses for Windows paths. It's also straightforward - just encode the unmatched surrogate pairs as if they were the corresponding reserved unicode characters using the normal UTF-8 algorithm.
- nyanpasu64 5y agoRust uses WTF-8 on Windows for OsStr[ing] and Path[Buf]. It's zero-overhead to cast from &str to &OsStr/&Path to &[u8] (though converting WTF-8 to UTF-16 costs an extra operation when performing a Win32 function call). However this doesn't solve the inability to round-trip "possibly-valid UTF-8/16" to "Unicode text" and back (though Python's surrogateescape might be one viable approach). Other libraries handle this even worse than Rust. On Linux (filenames are bytes), Qt is unable to open files with invalid UTF-8 names, while GTK can open them (but shows an "invalid encoding" message instead of the original filename), which I think is a good-enough approach.
- breakingcups 5y agoIt's also a common thing that Silent (aka CookiePLMonster) fixes in the games he patches. See for example: - https://cookieplmonster.github.io/2020/05/23/silentpatch-mafia-ii-definitive-edition/ https://cookieplmonster.github.io/2020/05/23/silentpatch-maf... - https://cookieplmonster.github.io/2021/02/27/silentpatch-yakuza-remastered-collection-r2/#part-2--silentpatch-for-yakuza-5 https://cookieplmonster.github.io/2021/02/27/silentpatch-yak...
- garaetjjte 5y agoPart of the problem is legacy Windows cruft. For long time to properly handle Unicode characers you needed to explictly use widechar UTF-16 functions. Legacy narrow encoding is systemwide setting, couldn't be set to UTF8, thus only subset of characters would be represented correctly. Only recently they introduced ability to set narrow encoding for application to UTF-8 with setlocale, which is a lot saner.
- mkotowski 5y agoIn case of a home-grown code, it could be simply the question of a programmer awareness. There are still many outdated and/or unfinished tutorials that use WinAPI without any concern about enabling Unicode and wide chars support. If we are talking about ready game engines like Unity and Unreal... it is probably a naive assumption about input being 1 byte wide and things getting lost because of that in some gamedev-made script.
- jan_Inkepa 5y agoI've been bitten on a few small releases by forgetting that C# localises number->string conversion by default (which makes sense. But if you forget, and you're writing floats to csv files and the decimal points become decimal commas....).
- account42 5y agoI disagree that having localization for number formatting based on a system setting by default makes sense. Formatted numbers are needed for both human and machine consumption and only one of those can deal with unexpected formatting.
- jan_Inkepa 5y agoMaybe the galaxy-brain design principle is: if you're designing an API, make sure that where possible bugs occur in an area where programmers care about fixing them (data I/O) rather than somewhere that they neglect (user interface localisation). Voila: better software!
- account42 5y agoExcept programmers test with their own locale and everything works there. Then the user gets an obscure error that the programmer is not able to reproduce because on their system a number from some internal config file was parsed incorrectly.
- GoblinSlayer 5y agoIt's text encoding confusion: https://en.wikipedia.org/wiki/Mojibake https://en.wikipedia.org/wiki/Mojibake