3 ms·
I believe that library / service is called UTF-8. These days everything should be stored as bare UTF-8 data (or utf8mb4 if you're MySQL) and presented without
by web007 5y ago
I believe that library / service is called UTF-8.
These days everything should be stored as bare UTF-8 data (or utf8mb4 if you're MySQL) and presented without anything else. Don't parse it, don't slice-and-dice it, don't prepend or append titles or honorifics or suffixes, don't make assumptions about length or content beyond "must be > 0 as a whole" and DEFINITELY don't use it as an identifier. Treat it as a non-unique opaque token and you'll be fine greater than 99% of the time.
There are people with no last name. There are people with two or three or twelve middle names. There are people with a number for a last name. There are people with a symbol for their entire name.
Take what they give you and use it and be done with it.
- lmm 5y agoNot good enough, thanks to Han unification - if you do that you'll mangle Japanese names.
- indigo945 5y ago"Mangle" is an exaggeration. Japanese names will look correct after Han unification both to Japanese people on a Japanese computer and to Chinese people on a Chinese computer. (All other combinations fail.)
- lmm 5y ago> Japanese names will look correct after Han unification both to Japanese people on a Japanese computer and to Chinese people on a Chinese computer. Only if you are displaying them in a way that respects the computer's preferences (most websites and programs, especially American websites and programs, don't) and those preferences are set correctly. And certainly if you have text blocks that contain both Chinese and Japanese names you will always mangle at least one of them.
- jack1243star 5y agoAnd we bilinguals can spam <span lang="..."> to try force the correct font on our blogs. Ugh.
- web007 5y agoTIL, https://en.wikipedia.org/wiki/Han_unification#Examples_of_language-dependent_glyphs https://en.wikipedia.org/wiki/Han_unification#Examples_of_la... It looks like there's no general solution possible with Han unification. If you have any two of ZH and JA and KO and VI in a page, you will fail to display one of them correctly for certain characters unless (as in that wiki page) you add a LANG attribute for each element they are contained within. Personally, I would use the browser's language or user locale to set the page language and give up. Then in Japan the local (Japanese) names would look fine, and same for China, Korea and Vietnam. Local consistency versus complicated perfection (tracking the input language as well as tokens and using them everywhere), and I could blame the browser for doing poorly at its impossible job. One possible "perfect" fix would be to store the token and <span lang=...>$token</span> as well. The only place the non-wrapped version would be used is plaintext email or SMS, either of which are beyond lost causes for other reasons. Doing it with an embedded SPAN tag presents its own problems with sanitization, as well as guaranteeing it's always wrong if the input language was specified incorrectly when the token was populated, versus as above where it would be corrected to the local-optimal version if the user locale overrides it.
- sushsjsuauahab 5y agoI pray that a more complicated solution is not needed, but when I was living abroad I would always encounter issues with sites where they thought "since the IP is from x, or since the browser is requesting lang y, then we should think that this American passport holder is a Spanish citizen and thus we can make assumptions about him." The ultimate source of this issue is that we are taking names and official IDs too seriously, but I doubt that problem will go away for "serious business". Funnily enough though, it already has for things like restaurant table reservations where all info provided is quite literally just a string for a human to do something with. No need to validate if the user's phone country code matches the country in which they are reserving a table...
- WorldMaker 5y agoIt's not an ideal fix, but Unicode has had "variation selectors" for a few versions now which can force specific CJK variations: https://en.wikipedia.org/wiki/Variation_Selectors_%28Unicode_block%29 https://en.wikipedia.org/wiki/Variation_Selectors_%28Unicode... Variation selectors are getting a good workout/testing technically in emoji at least (a lot of emoji are "just" "old" Unicode codepoints with a ZWJ and the variation selector known as the emoji variation selector to tell systems to always show it in "emoji styles"). I can't speak for how well it works in practice for CJK languages as I don't know them (more reason I appreciate emoji for letting me test compatibility with hard parts of UTF-8 in ways that I can read and most users want), but I do appreciate that there's at least the idea for/part of a fix in "recent" Unicode. I'm also imagining it is not a fun thing to implement in practice, as Unicode at this point maintains a massive database just for it: https://www.unicode.org/ivd/ https://www.unicode.org/ivd/
- GoblinSlayer 5y agoAmerican honorifics get me every time. What's the point in teaching a computer to use honorifics? It's a heap of semiconductors that stirs a heap of bits. And on top of that you teach it yourself.