5 ms·
Evolving how? I'm not aware of any reason to move beyond UTF8 for encoding Unicode.
by roca 5y ago
Evolving how? I'm not aware of any reason to move beyond UTF8 for encoding Unicode.
- _vvhw 5y agoWhich version of UTF8?
- coldtea 5y agoWhen adopting a new niche language, with hardly any following, and frequent changes, and not even an 1.0, like Zig, "which version of UTF8" (as if that's an issue) is the least of your worries... "Which third-party strings lib of several half-complete incompatible libs" will be a much realer concern...
- _vvhw 5y agoZig's C interop is pretty good though, and there must be some decent native Unicode library out there somewhere right? ;) I've worked on a full-duplex file synchronization system that had to support cross-platform operating systems and file systems, across a variety of Unicode normalization schemes (and versions [1], which is why I introduced the question), and I'm personally satisfied that baking this into the language specification would be a mistake. [1] For example, depending on the file system, there's simply no way to get the normalization right unless you reverse engineer the actual table they're using, or probe the file system to do the normalization for you.
- motiejus 5y agoWould Bellard's libunicode[1] work? [1]: https://github.com/bellard/quickjs/blob/master/libunicode.h https://github.com/bellard/quickjs/blob/master/libunicode.h
- jstimpfle 5y agoFor systems programmers the answer to "which third-party strings lib" is probably "None, write your own that fits with the rest of the system". A ready-made lib will be a lot of work to fit in - consider choice of internal encoding, allocation, hashing, buffering, mutable operations, etc.... Assuming that you really want to use UTF-8 internally, which is probably a sensible choice, the reusable part of a string library is basically the UTF-8 encoded/decoder. A useful implementation of UTF-8 is about 100-200 lines, I could probably rewrite what I use in an hour or two without an internet connection. The rest of the work is integration stuff that doesn't make sense to put in a library IMO. The idea of a string library fits much better with garbage collected and scripting languages (which includes C++ with RAII mechanism, but consider that std::string and similar often cause bad performance). Many programs, in particular non-graphical programs don't need any UTF-8 code at all - UTF-8 handling is basically memcpy().
- rightbyte 5y ago> Many programs, in particular non-graphical programs don't need any UTF-8 code at all - UTF-8 handling is basically memcpy(). argv to main is utf8 on my system.
- anonymoushn 5y ago> argv to main is utf8 on my system. That sounds totally compatible with programs that don't know anything about utf8. Do programs need to normalize the utf8 you pass in before using it as an argument to open(2) or something?
- jstimpfle 5y agoOn Unix/Linux, it is binary data without any restrictions except that each argument is zero-terminated (the typical argument is probably UTF-8 if you have set a UTF-8 locale). You'll see exactly the bytes that were put in as arguments to execlp() et. al. by the parent process. On Windows, I believe it is Unicode converted to current codepage. In any case I don't need to care about it since I can simply treat arguments as ASCII-extended opaque strings as described.
- chaz6 5y agoHow feasible would it be to defer string processing to the operating system so that the behavior of all software running on it is the same? Perhaps a new OS interface could be defined for this purpose using syscalls on Linux. At the very least, there should be one canonical set of algorithms per operating system, rather than everyone downstream reinventing the wheel. Please forgive me if this sounds absurd, I am not a low-level programmer.
- anonymoushn 5y agoI certainly don't want to pay for a syscall to do string encoding. Also, a lot of the time the problem isn't that people are using fundamentally incompatible string libraries, but that there isn't one correct answer to the question they're asking and they chose different ways to convert the question into code. A reasonable question to ask is "How many extended grapheme clusters are in this string?" The answer is "It depends on what font you plan to use to render it." Not great! Some programmers would still like to e.g. write a reverse() function that returns ":regional_indicator_f::regional_indicator_r:" unmodified (because it is the French flag emoji) and returns ":regional_indicator_i::regional_indicator_h:" when given ":regional_indicator_h::regional_indicator_i:". If such people want to avoid having nonsensical behavior in their programs, the only solution available is to decide what domain the program should work on and actually deal with the complexity of the domain chosen.
- ShrigmaMale 5y agoBecause string processing is not slow enough? Making string operations eat the performance impact of that context switch in and out of kernel is not a good idea. Library is better. This is not string but I'm thinking to some language like Rust where the "time" crate is not language feature but just about standard and all the other library use it. This is possible if a library like it is good quality and exist early in the language.
- roca 5y agoRegular UTF8, not WTF-8 or any of those other variants (which are for encoding data that is not necessarily Unicode).
- _vvhw 5y agoAlso excluding Unicode normalization? Or should that also be baked in?
- roca 5y agoNo need to drag Unicode normalization into it; don't require strings to be normalized. Normalization is only relevant in very specific contexts and you don't want to pay for it elsewhere.
- _vvhw 5y agoAgreed, but I think that many people would consider Unicode normalization to be part of what they want from the std lib when they mean that UTF8 should be baked in... so that they can manipulate UTF8 as they want, including in various normal forms according to platform. It's hard to imagine people being satisfied without having access to Unicode normalization. For example, consider JS' introduction of String.normalize(). This is a slippery slope. It had a huge impact on Node's build process and binary sizes because now all the tables had to be shipped. But it's still broken in JS, because no matter the Unicode normalization support provided, it will never match the exact tables used e.g. in Apple's HFS. I feel that by the time it gets to String.normalize(), it's too far gone.