4 ms·
UTF support?
by Keyframe 7y ago
UTF support?
- jerf 7y agoAs much as C strings support it.
- rurban 7y agoExactly. You should not name a buffer lib "string", when it does not support the basic unicode operations: case fold, normalize => compare, search. In utf-8 of course. I'm also missing stack allocation support, needed for fast short strings. It should be even included in sdsnew, for len < 128.
- Keyframe 7y agoUnfortunately, yes. Limited usefulness at best without unicode support, at least to a degree. Even UTF-16 or 32 internally would suffice, treating UTF-8 only as ser/de format is good enough these days.
- jstimpfle 7y agoUTF-8 is the preferred internal storage format for most applications. The reason is space efficiency.
- jandrese 7y agoThat and the perceived simplicity of UTF16/32 is not actually the case so why bother if you have to do it the hard way anyway.
- cardiffspaceman 7y agoI certainly have leaned on UTF-8 but I wonder how efficient it is for the numerically-higher code points if the language is not heavily cp1252, like Korean.
- pstch 7y agoAn interesting - but not surprising - thing about this is that compression algorithms can be more efficient on wider representations of numerically-high code points (e.g, for some Korean corpus, using UTF-32 instead of UTF-8 improves LZMA compression by ~10%).
- zzo38computer 7y agoHow well does that corpus compress with LZMA if using a Korean specific character code (such as EUC-KR)? And what about other combinations, with other character codings and other compression algorithms?
- pstch 7y agoEUC-KR doesn't improve much with LZMA (2% over UTF-16), but is better with gzip-9 (10% over UTF-16). I haven't studied this extensively, just did a few tests when waiting for it to download.
- asveikau 7y agoI don't get entirely why people want or expect to store UTF-32 in memory for any substantial period. It makes more sense to me as: if you need to process codepoint at a time, parse from UTF-8 or UTF-16 in a codepoint-at-a-time fashion, leaving only ~1 codepoint decoded at a time. People seem to think that 1 codepoint = 1 integer frees you from thinking Unicode is hard. But 1 codepoint is not 1 glyph. You have combining characters, zero-width joiners [used also in emojis], RTL markers, Han unification, probably more. So you can't really think of a Unicode string as a random-access, one-glyph-per-unit type of thing in any encoding. So I hope when people say "UTF-32 support" they mean "decode UTF-8".
- zzo38computer 7y agoThe Unicode functions of Glk use UTF-32 (and numbers outside of Unicode range are used for function keys), and Glulx supports strings of 32-bit characters (usually Unicode, although this is not required) (strings of 8-bit characters are also supported, but this doesn't support UTF-8 or any other multibyte encoding). (However, strings in Glulx are usually huffed anyways, so other than the temporary buffers you would ordinarily store a variable number of bits per character (or sometimes a string of several characters has its own code), and can be shorter than using UTF-8 or UTF-32 or other uncompressed encodings.) The stuff you mention does make Unicode very messy though.
- unwind 7y agoHow do you implement stack allocation support inside a C library function? How is de-allocation expressed?
- ChickeNES 7y agoalloca()?
- unwind 7y agoNo. Quoting the manual page: The alloca() function allocates size bytes of space in the stack frame of the caller. This temporary space is automatically freed when the function that called alloca() returns to its caller. That would not work, since the code calling alloca() is the library, and you want the to allocate on the stack frame of the function calling the library. I don't think that is possible, unless you make the library function into a macro (perhaps using ?: on the allocation size).
- noobermin 7y agoFor stack allocation might as well just make a fixed sized char array then. sds by its design seems meant for heap strings (d in the name means dynamic).
- faragon 7y agoFor C string library with stack allocation and some Unicode support you can check: https://github.com/faragon/libsrt https://github.com/faragon/libsrt
- Vordax 7y agoShould just use the ICU string library. You get utf support and much more that is needed for modern string usage. https://unicode-org.github.io/icu-docs/apidoc/released/icu4c/ https://unicode-org.github.io/icu-docs/apidoc/released/icu4c...