4 ms·
>> The Winapi (aka Win32) is the only commonly used API that regularly requires something other than UTF-8 (the Windows Unicode APIs use UTF-16 – not UCS-2 as i
by mau 13y ago
>> The Winapi (aka Win32) is the only commonly used API that regularly requires something other than UTF-8 (the Windows Unicode APIs use UTF-16 – not UCS-2 as indicated in the Spolsky article).
> Why has Microsoft still not fixed this?
Is there a way they can fix it without breaking backward compatibility?
- ygra 13y agoThey can, by introducing a few thousand stub functions which convert and delegate to ★W functions, but there'd be little use to do so. Every sane program out there uses the ★W functions and some insane still use the ★A ones. So it would only be beneficial for new code while all existing code remains the same, with the same encoding bugs if there are any. I'm also not sure whether there are that many cases where it really helps. UTF-8 only on Windows is painful and so is using UTF-8 only with all other things that use UTF-16 (Qt, Java, etc.). Usually in those cases you use a library/framework/whatever that handles the platform abstraction or just conform to what's expected.
- AnthonyMouse 13y ago> Every sane program out there uses the ★W functions and some insane still use the ★A ones. So it would only be beneficial for new code while all existing code remains the same, with the same encoding bugs if there are any. Every sane program uses the undifferentiated functions, defines UNICODE so that they map to the ★W functions, and uses TCHAR which defining UNICODE causes to map to a wide char. Older programs don't define UNICODE and often use char (or CHAR) instead of TCHAR. If they would create a different define (e.g. '#define UTF8') which would map to the new UTF8 functions and would define TCHAR as CHAR then anything doing it either way would do the right thing just by defining UTF8 and recompiling. Only programs that explicitly call the ★W functions (which they never should have exposed) wouldn't be "fixed" to use UTF8, but neither would they be broken. > UTF-8 only on Windows is painful ...because Microsoft hasn't fixed it. > Usually in those cases you use a library/framework/whatever that handles the platform abstraction or just conform to what's expected. That's a cop out. You're just deferring to the frameworks, which also shouldn't be using anything other than UTF8, and who may have more difficulty in fixing it because the transition mechanism Microsoft used to unicode is well adaptable to another transition. Not every library you have to use will use the same encoding as the framework and you're back to a huge pain. The only way to fix it is for everything to always use UTF8, and deprecate everything else going forward.