4 ms·
UTF-8 encoding works "as is" based on byte strings (char[]). The latest versions of the draft standard provide somewhat more support. I recommend heading towa
by DougGwyn 6y ago
UTF-8 encoding works "as is" based on byte strings (char[]). The latest versions of the draft standard provide somewhat more support.
I recommend heading toward a future where only UTF-8 encoding is used for multibyte characters and UCS-2 or similar for wchar_t. There is no need to support several different encodings.
- deleted 6y ago[deleted]
- rseacord 6y agoAaron Ballman even got a u8 character prefix added to C2x: N2198 2018/01/02 Ballman, Adding the u8 character prefix http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2198.pdf http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2198.pdf
- ori_b 6y agoUCS-2 is a bad choice -- it fails to represent most unicode characters. If you meant UTF-16, that's also a bad choice, because UTF-16 is also a variable width encoding, forcing programmers to use a some for of "extra-wide char". I'm of the opinion that wchar_t should become an alias for char32_t.
- DougGwyn 6y agoYes, I meant the 31-bit code point value (more than 16, anyway). It is the most useful width for doing things with wide characters.
- a1369209993 6y agoUTF-32 is also a variable-width encoding; eg 00000044 00000308 aka "D̈".
- DougGwyn 6y agoI thought it was strictly one character per 32-bit code. Anyway, whatever it is called it is what wchar_t should be.
- a1369209993 6y agoThere are no fixed width encodings with range of encodable characters anywhere near that of Unicode.
- flatfinger 6y agoIt's too bad Unicode wasn't designed around the concept of easily-recognizable grapheme clusters and "write-only" [non-round-trip] forms that are normalized in various ways. A text layout engine shouldn't have to have detailed knowledge of rules that are constantly subject to change, but if there were a standard representation for a Unicode string where all grapheme clusters are marked and everything is listed in left-to-right order, and an OS function was available to convert a Unicode string into such a form, a text-layout using that OS routine would be able to accommodate future additions to the character set and and glyph-joining rules without having to know anything about them.
- a1369209993 6y agoYou can't do that without commiting to not supporting pathological text, otherwise you're stuck adding new special cases to the layout engine every update anyway. I do have some ideas for a better encoding (like, I assume, anyone competent with sufficient free time and interest in text encoding), but there's a lot of reluctance to put effort into something that's already completely eclipsed by a technically inferior but not completely unusable alternative, so I've had it mostly shelved.