3 ms·
I think I'll disagree, and I'm someone with a native Asian language. A character doesn't mean anything in Asian languages, and any attempt to use a fixed-length
by lopsidedBrain 7y ago
I think I'll disagree, and I'm someone with a native Asian language. A character doesn't mean anything in Asian languages, and any attempt to use a fixed-length encoding is pointless. The concepts you want instead are either a code point or a glyph. The concept of a code point is useful as a part (and not the whole) of Unicode-validating, encoding, and decoding. A glyph is useful mostly in rendering engines (i.e. webkit-internals and UI rendering frameworks). Okay, maybe they are also kind of useful for sorting and collating. But fixed-width character encodings are almost never useful, and invite programmer to make assumptions about how strings can be sliced.
Most server-side applications should never have to know what these concepts even are. Or any library that is not user-facing. They get bytes from the UI layer, and they can keep them as opaque bytes. For user-facing apps, you can ask your renderer library for a pixel-width or similar for a string, and let them handle how to parse it. Very little code ever needs to know about unicode.
Any kind of input-sanitization is vastly simplified in utf8, and that makes it worth it for me. For me the really troubling trends are conventions like Rust Utf8Error, where they can cause what I'd consider a UI-related exception in code that had no business even interpreting what those bytes are. Unfortunately, every API uses strings, so they are kind of hard to avoid. It introduces what I'd consider a software layering problem.
Maybe others here with more experience with internationalization can chime in and tell me I'm wrong.
- Freak_NL 7y agoAlso, UTF-16 is not fixed width in modern usage. As soon as someone uses an emoji, boom!, surrogate pairs. So unless UTF-32 is used at a glorious four bytes per character, you won't get fixed width. > For me the really troubling trends are conventions like Rust Utf8Error,[…] Interesting. Isn't that only returned when the input bytes contain a non-UTF-8 byte sequence? How is Rust's approach different from other languages?
- atoav 7y ago> Interesting. Isn't that only returned when the input bytes contain a non-UTF-8 byte sequence? As a Rust user: this is what your function returns if the user inputs non-UTF-8 bytes into something that expects UTF-8 bytes and the programmer explicitly choses not to handle that error. I don’t see anything wrong with that. Sometimes you might be interested in receiving, processing, storing valid UTF-8 strings rather than arbitrary byte sequences, that may or may not be able to be translated back into something valid that you can display. I always hated to do encoding related work with a passion before I started using Rust. Rust forced me to do it the right way and actually understand why I am doing it a certain way. When it comes to encoding I feel safer in Rust then e.g. in Python, despite having used it for four times as long.
- camgunz 7y agoUsing any UTF encoding as fixed-width is almost always a mistake. There's no concept of "characters" in Unicode or the UTF encodings, they use codepoints, and those are only very rarely useful--you cannot treat them as "characters". This is a common misconception with UTF encodings.
- camgunz 7y agoI think we actually agree. > A character doesn't mean anything in Asian languages, and any attempt to use a fixed-length encoding is pointless. Totally agree, indexing into a string is bad practice, no matter the encoding (because you probably can't ever guarantee what the encoding will be, or how it was converted, etc.) This is true for UTF-8 and UTF-16, and really any encoding because again, you can't be sure what you're dealing with. Code points, "characters", glyphs, etc. are all concepts that work at different parts of the stack (well, characters doesn't), which again is true for all encodings. The advantage UTF-16 has for languages that tend to be multibyte is it's representation takes up less space. Other than that, it has all the same disadvantages any other encoding has. > Any kind of input-sanitization is vastly simplified in utf8, and that makes it worth it for me. For me the really troubling trends are conventions like Rust Utf8Error, where they can cause what I'd consider a UI-related exception in code that had no business even interpreting what those bytes are. Unfortunately, every API uses strings, so they are kind of hard to avoid. It introduces what I'd consider a software layering problem. This is the problem right here. The "UTF-8" the world initiative ignores that UTF-16 is a lot more practical for most people, and as a result you get platforms expecting UTF-8 that really have no business doing so.