5 ms·
Trying to handle character encoding on Windows in multi-platform programs is a nightmare. In C++ you can almost always get away with treating C strings as UTF-8
by alexbock 10y ago
Trying to handle character encoding on Windows in multi-platform programs is a nightmare. In C++ you can almost always get away with treating C strings as UTF-8 for input/output and you only need special consideration for the encoding if you want to do language-based tasks like converting to lowercase or measuring the "display width" of a string. Not on Windows. Whether or not you define the magical UNICODE macro, Windows will fail to open UTF-8 encoded filenames using standard C library functions. You have to use non-standard wchar overloads or use the Windows API. That is to say, there is no standard-conformant internationalization-friendly way to open a file by name on Windows in C or C++. I really wish Microsoft would at least support UTF-8, even if they want to stick with UTF-16 internally.
The section titled "How to do text on Windows" on http://utf8everywhere.org/#windows http://utf8everywhere.org/#windows covers the insanity in more detail.
- Ono-Sendai 10y agoIt's a bit annoying and lame but not really a huge deal. For example, you can just open an fstream like so: std::ofstream file(convertUTF8ToFStreamPath(pathname).c_str()); Where convertUTF8ToFStreamPath converts from UTF-8 to the Windows wide encoding (UTF-16?)
- alexbock 10y agoIt gets even more fun when you have no choice but to use a poorly maintained closed-source third-party library (par for the course on Windows). If the library authors didn't use this trick themselves you can get stuck in situations where you have no way to make the library open the file without resorting to the DOS 8.3 name. And if the user disabled 8.3 names system-wide... then the only consolation is that something even more important on the system will probably break first. A sad state of affairs in 2016.
- Ono-Sendai 10y agoThe third party library situation is indeed trickier. I've never found a library that I couldn't get to open Unicode paths eventually though. Sometimes you have to open a file handle yourself and pass it to the library. Edit: By the way the worst handling of Unicode paths on Windows I have found is by Ruby, which is still partially broken. (last time I checked)
- cremno 10y agoWhat is broken? Have you reported it? I've reported a related bug this week (File.truncate called CreateFileA() with a UTF-8 string). I even had a working patch but forgot to attach it. Anyway it was almost immediately fixed including various additional tests for other File class methods (which handled Unicode without any problems).
- Ono-Sendai 10y agohttps://bugs.ruby-lang.org/issues/1685 https://bugs.ruby-lang.org/issues/1685 etc..
- CountSessine 10y agoFor a company that claims to be so supportive of "developers, developers, developers", Microsoft's stubborn and developer-hostile approach to internationalization and their dogged loyalty to the awful UTF-16 encoding is ironic. The Right Thing To Do at this point is to make UTF-8 a multi-byte code page in Windows and build a UTF-8 implementation in the msvc libc. The milquetoast excuse I hear from Microsoft people is that some win32 APIs can't handle MBCS encodings with more than 3 bytes per character. Which sort of sounds like a problem for developers to fix; perhaps Microsoft could hire some?
- adzm 10y agoEspecially hilarious how SQL server nchar is widechar so most all string data is twice as large as it needs to be. But there isn't a utf8 option to enable! Though this might have been addressed in one of the later versions, I wouldn't be surprised if it was limited to enterprise edition or other crap.
- chris_wot 10y agoBloody hell... I looked this up as I was certain it couldn't still be the case. But yes, it is! This was logged as an issue on Connect in 2008, [1] and Microsoft's response was: "Thanks for your suggestion. We are considering adding support for UTF8 in the next version of SQL Server. It is not clear at this point if it will be a new type or integrate it with existing types. We understand the pain in terms of integrating with UTF8 data and we are looking at ways to effectively resolve it." 1. https://connect.microsoft.com/SQLServer/feedback/details/362867/add-support-for-storing-utf-8-natively-in-sql-server https://connect.microsoft.com/SQLServer/feedback/details/362...
- bitwize 10y agoBefore "developers, developers, developers" comes "backward compatibility, backward compatibility, backward compatibility". Windows is perhaps the first commercial platform to commit to Unicode; they made that commitment when UTF-8 was still some notes scribbled on Brian Kernighan's napkin. And all future Win32 implementations must be 100% binary compatible with previous ones. That creates inertia for UTF-16 (or UCS-2), true, but the backwards compatibility guarantees make Windows an absolute joy compared to Linux if you want to write software with a long service lifetime. The decision to stick with 16-bit Unicode is an engineering tradeoff.