10 ms·
The Wonderfully Terrible World of C and C++ Text Encoding APIs (With Some Rust)
- mlindner 4y agoI'm personally a big fan of saying "we only support UTF-8, if you feed us UTF-16 please contact the creator of the software that uses UTF-16 to fix their use".
- int_19h 4y agoIt's kind of amazing that something as basic as character encodings - at least the basics like UTF-16 ↔ UTF-8 ↔ stdio-encoding! - is something that's still not in the C++ standard library. For a while there was codecvt_utf8 et al, but that was deprecated 5 years ago in C++17 with no replacement "to clear the path for the future" (https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2017/p0618r0.html https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2017/p06...), yet no replacement came in C++20, and none are planned for C++23.
- kevin_thibedeau 4y agoUnicode support requires incorporating their database into a library. At a minimum you need to know which code points are combining chars. For a language with five to ten year update cycles should everyone be stuck with outdated data if the Unicode standard is revised in the interim?
- Ferrotin 4y agoConversion between encodings doesn’t require a database or knowledge of combining characters.
- poorlyknit 4y agoThis. But it strenghtens the arguments that programming environments should just come with some sort of support for the most common encoding forms.
- nine_k 4y agoIt does: say, Latin-1 has one-byte code points for characters like ê, but a UTF-8 sequence may represent it as an e followed by a combining circumflex.
- lultimouomo 4y agoLatin-1 is a different character set; you do not need knowledge of code points to convert between different UTF encodings.
- int_19h 4y agoThat does not require you to know which Unicode characters are classified as combining marks. It only requires that a Unicode sequence U+0045 U+0302 is translated to Latin-1 character 0xEA (or vice versa).
- kevin_thibedeau 4y agoA library that only extracts code points will do more damage than not having one at all. If you have to decode Unicode you presumably want to parse it some of the time. Not supporting the needs for string processing with multi-point graphemes leads to broken Unicode "support" that doesn't actually work with all valid Unicode.
- Ferrotin 4y agoPlenty of applications just need to take strings in one format and pass them along in another format without doing feats of string processing to them.
- int_19h 4y agoTo elaborate: UTF-8 is the most common I/O encoding, but UTF-16 is often what you need to work with the OS APIs (Win32, macOS) or popular frameworks (Qt).
- TillE 4y agoThe funny thing is that C++ sort of has a flexible string class which uses system-dependent conversions, std::filesystem::path It's super convenient for working with Windows file APIs, but everything else is still a pain.
- Ferrotin 4y agoYep. There are still some annoyances like invalid UTF-16, so maybe there is more is needed than just the basic transformations — some precise handling or parameters specifying how to handle these edge cases could be necessary. Handling UTF-8 versus WTF-8, and properly round-tripping these representations, is probably a bigger issue than graphemes and canonical forms for most users.
- josephg 4y agoAlso unfortunately the web browser. Javascript strings are UTF-16 (Well, UCS2).
- arka2147483647 4y agoAll operating systems have the unicode database saved somewhere. There should just be a standard way of accessing it. Just like filesystem. Edit; that is; it does not have to be linked in the standard lib. Can be a data file somewhere, or a a shared lib.
- duskwuff 4y agoThere's precedent for this, too: time zones! Time zone data can change over time, and as such it's typically stored in system files and loaded at runtime, rather than being embedded in executables. Locales have some similar behavior as well.
- thwarted 4y agoLike the timezone database, many things of which ship(ped) with their own copy.
- apaprocki 4y agoIt's funny how much time I (and probably a number of others) spend undoing this decision and making things all refer to one copy of the data. There's usually also one per programming language/runtime, so the list only grows as new languages/runtime keep coming into the picture. Everyone's so good at patching vulnerabilities, but did you plan for your app breaking in production for all users in particular timezones because a random service in a chain of a 6, say, re-implemented in Go last quarter, or maybe... changed their Node app to port from using Moment to the standard Intl object 10 months ago, thus depending on the ICU C++ library, and a separate copy of the data that is no longer correct? Try to assess that risk and your head will hurt. These are the sharp edges in a polyglot environment that usually aren't considered when deciding to bring in new languages. Unicode CLDR has the same issue, but the pain usually isn't acute because it takes much, much longer for new glyphs or other data to appear in common use or start appearing in data feeds or from OS APIs/input. As long as you're at least on some update schedule, it's usually fine, and there's usually fewer things to update.
- duskwuff 4y ago
- lultimouomo 4y agoI feel your pain. Last week I just gave up and wrote my UTF-8 to UTF-32 conversion routine. It took me far less to do that than I spent looking for a standard solution.
- usrnm 4y agoConversion is the easy part, though. Now try to count the number of "characters" in a string. You will quickly realise that it's much easier to just reach for icu, or at least some parts of it
- mlindner 4y agoThat's not something you should be doing anyway. It's not a useful thing to talk about and doesn't even work for all languages. If you are writing some software that requires counting the number of characters in a string you should stop and rethink why you're trying to do what you're doing.
- lultimouomo 4y agoSure, but a) conversion was exactly what was provided but C++11, before they deprecated it b) plenty of staff that the standard library does is complicated! Some of it also seems less fundamental than handling plain text.
- int_19h 4y agoIn practice, counting the number of "characters" turns out to be not all that common, though. This usually comes up doing low-level text rendering or editing, but how often would you do something like that yourself instead of delegating to some library? (and then that library might still reach for ICU - but that's another story)
- Findecanor 4y agoDoes it perform validation? Byte sequences too short or too long are obviously invalid, but codes encoded in too many bytes than necessary are too - which isn't immediately obvious.
- tialaramex 4y agoThe needs to convert character encodings should be gradually going away. There's not going to be more call for a standard library to address this in 2026 than there is today. It's more crazy that C++ was so slow to mandate UTF-8 support. You have the situation where your modern C++ environment might have several distinct "string" types none of which is guaranteed to just be UTF-8.
- vvanders 4y agoA lot of the contexts where C++ has been deployed predate the shift to UTF-8 so I think there's actually more of a reason to have the support.
- nine_k 4y agoSay, for Japanese texts Shift-JIS is vastly more economical than UTF-8, so a number of systems uses it. Likely this is the case for some other writing systems. UTF-16 should work in there cases eventually though.
- astrange 4y agoMany of those texts include enough ASCII characters (like HTML tags) that UTF-8 is more efficient even then.
- tialaramex 4y agoFor ASCII (except the blackslash and tilde) Shift-JIS and UTF-8 are the same, one byte. So your HTML tags work just fine in Shift-JIS. For small Japanese-only text, Shift-JIS really will be smaller, but it's just not usually enough to care, and the price is you can't do Unicode. It occupies a similar space to 8859-1 / Windows 1252 where older software uses it but you should just transition.
- astrange 4y agoOh, I meant UTF-8 would be more efficient than UTF-16.
- jcelerier 4y ago> It's kind of amazing that something as basic as character encodings - at least the basics like UTF-16 ↔ UTF-8 ↔ stdio-encoding! - is something that's still not in the C++ standard library. how would that work if you're on a microcontroller ? if $GOV_AGENCY says "product XXX is made according to international standards, including ISO/IEC 14882:2026" and international standard ISO/IEC 14882:2026 now says that a complete Unicode implementation has to be provided otherwise you're not compliant, any device with less than the 20-30-ish megabytes of memory needed for the unicode database won't be able to pass certification even if they don't handle text at any point, e.g. it's some dsp filter somewhere in a camera lense
- klodolph 4y agoVarious parts of the standard are marked as OPTIONAL, so you do not need to support them to have a compliant implementation. This includes, at the moment, any kind of Unicode. You may be aware that C and C++ are still used in legacy environments... "legacy" in the sense that these environments have been in continuous use for a long time. > __STDC_ISO_10646__ > An integer literal of the form yyyymmL (for example, 199712L). If this symbol is defined, then every character in the Unicode required set, when stored in an object of type wchar_t, has the same value as the code point of that character. The Unicode required set consists of all the characters that are defined by ISO/IEC 10646, along with all amendments and technical corrigenda as of the specified year and month. You can see the wording, "if this symbol is defined". Note that even if your implementation supports Unicode, and has all sorts of character conversion tables and character property tables, you can easily arrange for those to be statically linked, and only included if actually used. This is normally how standard libraries work in programming environments for resource-constrained systems like microcontrollers. Another interesting macro is the __STDC_HOSTED__ macro--basically, if __STDC_HOSTED__ is missing, then large chunks of the library will not be available. Still completely standards-compliant.
- Hello71 4y agoyou only need 30 MB if you use the ridiculously bloated icu library. musl implements POSIX locale APIs including iconv and wcwidth plus the entire rest of the libc in around half a megabyte. if you need case folding, character classes, or other special functionality, utf8proc is about 0.3 MB, and libunistring is about 1.5 MB. additionally, if your device has less than a few megabytes of storage, you'll almost certainly be using static linking, which automatically prunes unused code from properly designed libraries.
- svnpenn 4y ago> ICU has almost everything about their APIs correct No. If you have a common function with 13 arguments: U_CAPI void ucnv_convertEx( UConverter *targetCnv, UConverter *sourceCnv, // converters describing the encodings char **target, const char *targetLimit, // destination const char **source, const char *sourceLimit, // source data UChar *pivotStart, UChar **pivotSource, UChar **pivotTarget, const UChar *pivotLimit, // pivot UBool reset, UBool flush, UErrorCode *pErrorCode); // error code out-parameter you've done something terribly wrong. Refactor it as a method, do "factory" stuff, something. Just shocking what C people are able to put up with.
- the_svd_doctor 4y agoA method in C?
- TillE 4y agoIt's an extremely common pattern in C to have a struct with a bunch of associated functions which behave exactly like methods, taking a pointer to the struct as their first argument.
- eyelidlessness 4y agoYup. This is exactly how Python methods work too. And how most functional code looks in any language.
- olliej 4y agoobject oriented code does not mean "language support for object oriented code". Object oriented code is very common in many larger C libraries (or things like the linux kernel). The only difference is that you don't get any compiler support to prevent errors. The standard model is: typedef struct __MyType MyTypeRef; struct MyTypeMethodTable { // the destructor - honestly at an api level you should probably have retain/release instead void (destroy)(MyTypeRef _this); // void (someMethod)(MyTypeRef _this, int someArg); }; struct __MyType { MyTypeMethodTable vtable; }; void MyType_destroy(MyTypeRef _this) { _this->vtable->destroy(_this); } void MyType_someMethod(MyTypeRef _this, int someArg) { _this->vtable->someMethod(_this, someArg); } Then an actual type is implemented as struct MyConcreteType { struct __MyType base; // some fields }; void MyConcreteType_destroy(MyTypeRef value) { MyConcreteType realValue = (MyConcreteType )value; // cleanup anything you need to do free(realValue); } void MyConcreteType_someMethod(MyTypeRef value, int someArg) { printf("I: %d\n", someArg); } MyTypeMethodTable MyConcreteType_MethodTable { .destroy = MyConcreteType_destroy, .someMethod = MyConcreteType_someMethod }; MyTypeRef CreateMyConcreteType() { MyConcreteType *result = calloc(1, sizeof(MyConcreteType)); result->base.vtable = & MyConcreteType_MethodTable; result->someField = whatever; return &result->base; // Or similar. avoid UB in C can make this weird } You can see that trivially this is pretty much what notionally OO languages like C++, Java, Haskell, etc do. An apparently not-uncommon error that happens in COM is: someObject->whateverTheirMethodTableIsCalled->someMethod(theWrongObject) Possibly with someObject and theWrongObject the other way around. The end result is sadness either way. OO programming is super effective for many things, and generally better for a lot of design, especially for libraries and frameworks. But nothing about OO requires compiler/language support - indeed the first discussions of OO code predated OO languages - but it's hopefully obvious to see that if the compiler can manage this, it results in less code, and less opportunity for error, and because the compiler manages those semantics it should technically produce better code (e.g if the compiler knows that "vtable" is actually a vtable pointer, it knows there is no circumstance it can change[1]). [1] Yes you could have incorrect code through UB, but in that case the compiler is allowed to do the "wrong" thing.
- nynx 4y agoAmazing how the C++ standardizing body makes APIs that are both way too generalized and yet not flexible enough to be highly usable and safe.
- olliej 4y agoThe Mac/iOS APIs are .. simple CFStringRef CFStringCreateWithBytes(CFAllocatorRef alloc, const UInt8 *bytes, CFIndex numBytes, CFStringEncoding encoding, Boolean isExternalRepresentation); CFDataRef CFStringCreateExternalRepresentation(CFAllocatorRef alloc, CFStringRef theString, CFStringEncoding encoding, UInt8 lossByte); Nice and easy, perhaps a teeny tiny bit inflexible? :D