2 ms·
So when writing Windows applications with Win32, would you use UTF-16 internally and read/write UTF-8 files or use UTF-8 internally and convert to/from UTF-16 a
by asifsignupfor 13y ago
So when writing Windows applications with Win32, would you use UTF-16 internally and read/write UTF-8 files or use UTF-8 internally and convert to/from UTF-16 at the boundaries to Win32 functions? My company currently does the first, UNICODE is defined, TCHAR=UTF-16.
- gilgoomesh 13y agoI would never use anything except UTF-8 for a general text file – for the efficiency reasons listed above – unless there was a strong reason otherwise. General text files are the biggest source of problems because they have no metadata indicating what encoding they actually contain. For everything else, it really depends what your data is for. If your data is only ever going to hold a file path that you need to pass to the Winapi, then there's no real problem with UTF-16. Although needing to have multiple paths for your text handling can become an issue. I write multi-platform C++ programs and even on Windows I use UTF-8 for all internal strings – including those that will eventually be passed to the Windows API. It just makes string handling simpler. However, I use a range of abstraction classes around all OS calls that transparently converts to/from UTF-16 as needed. A good example is boost::filesystem for paths and file I/O which internally stores UTF-16 on Windows but abstracts the need for me as the programmer to know or care about that detail – instead I can use UTF-8 everywhere and let the abstraction handle the encoding.