13 ms·
Your code displays Japanese wrong
- Asooka 5y agoCan't this be solved somewhat by adding a "cjk mode" zero-width character, like we have right-to-left/left-to-right embedding characters? Yes, yes, it's yet another standard, but there doesn't seem to be any way to indicate in the text stream itself what characters to use otherwise.
- flerovium 5y agoIt's illuminating to read this article's HTML: (Apologies, non-monospaced) However, this also meant that characters which differ in appearance across languages, such as <span xml:lang="ja" lang="ja">刃</span> and <span xml:lang="zh-Hans" lang="zh-Hans">刃</span> and <span xml:lang="zh-Hant" lang="zh-Hant">刃</span>, were given <strong>identical code points!</strong>
- madsohm 5y agoThe second character displays wrong for me when copying into VS code with Cascadia Code PL font.
- skhr0680 5y agoThat’s why Han unification is a mess.
- zzo38computer 5y agoMy own programs are specifically designed to not use Unicode. I think that Unicode is really messy and I dislike it. If you want to display Japanese text, EUC-JP can be used.
- Matheus28 5y agoI really hope you're being sarcastic
- zzo38computer 5y agoI am not sarcastic. I don't like Unicode.
- patrec 5y agoIf you like, to stick just to Japanese, this: https://upload.wikimedia.org/wikipedia/commons/b/ba/JIS_and_Shift-JIS_variants.svg https://upload.wikimedia.org/wikipedia/commons/b/ba/JIS_and_... better than unicode, you can't be helped. Not that I like unicode much either -- amongst other things the idiotic arrangement of codepoints makes it basically impossible to do remotely efficient text processing; e.g. here's a graph of the automaton the RE2 uses to check if something is an uppercase character: https://swtch.com/~rsc/regexp/cat_Lu.png https://swtch.com/~rsc/regexp/cat_Lu.png (For ascii there would exactly be a single arrow connecting two nodes).
- arp242 5y agoAnd here I was thinking that "Unix variants history" or "Linux audio systems" graphs were messy and complicated...
- yorwba 5y agoThis is about display, not encoding. Using EUC-JP to store text doesn't guarantee that it will be rendered with a Japanese font.
- lmm 5y agoIn practice it does, if it will be rendered at all. Elsewhere on this very page you can find people suggesting storing the display locale alongside the unicode string, which is really the only way to solve this problem in the general case - but in that case you might as well store pairs of byte sequence and encoding, there's not much difference between that and unicode string and locale.
- kalleboo 5y agoAren't the modern text display APIs of the most popular OSes all Unicode-based now? It seems likely that they will convert to Unicode when told to display a string in a different codepage and replace the locale info with the default Unicode behavior (of basing it on the user locale)
- yorwba 5y ago> In practice it does, if it will be rendered at all. Seems like you're right, at least as far as Firefox is concerned. Testing the data links below, it appears to guess the default language based on the encoding used. Neat! data:text/html;charset=euc-jp;base64,PHA+RGVmYXVsdDogv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9ImphIj5sYW5nPSJqYSIgv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9ImtvIj5sYW5nPSJrbyIgv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9InpoLWNuIj5sYW5nPSJ6aC1jbiIgv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9InpoLWhrIj5sYW5nPSJ6aC1oayIgv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9InpoLXR3Ij5sYW5nPSJ6aC10dyIgv8/EvrOks9G5/Mb+PC9wPgoK data:text/html;charset=euc-kr;base64,PHA+RGVmYXVsdDog7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9ImphIj5sYW5nPSJqYSIg7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9ImtvIj5sYW5nPSJrbyIg7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9InpoLWNuIj5sYW5nPSJ6aC1jbiIg7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9InpoLWhrIj5sYW5nPSJ6aC1oayIg7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9InpoLXR3Ij5sYW5nPSJ6aC10dyIg7NPywfqtysfN6ez9PC9wPgoK data:text/html;charset=gb2312;base64,PHA+RGVmYXVsdDogyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9ImphIj5sYW5nPSJqYSIgyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9ImtvIj5sYW5nPSJrbyIgyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9InpoLWNuIj5sYW5nPSJ6aC1jbiIgyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9InpoLWhrIj5sYW5nPSJ6aC1oayIgyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9InpoLXR3Ij5sYW5nPSJ6aC10dyIgyNDWsbqjvce5x8jrPC9wPgoK data:text/html;charset=big5;base64,PHA+RGVmYXVsdDogpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9ImphIj5sYW5nPSJqYSIgpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9ImtvIj5sYW5nPSJrbyIgpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9InpoLWNuIj5sYW5nPSJ6aC1jbiIgpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9InpoLWhrIj5sYW5nPSJ6aC1oayIgpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9InpoLXR3Ij5sYW5nPSJ6aC10dyIgpGKqva78qKSwqaRKPC9wPgoK
- innocenat 5y agoAnd what if you want to display Japanese and Korean at the same time?
- sleepy_keita 5y agoOh, this explanation is perfect to give to people when I encounter this error. Thanks!
- captainmuon 5y agoHow do you do it correctly in a bi-lingual app? Say your app is in English but you want to display asian language file names. Is there any way to tell if a string is chinese or japanese? I think CJK variation selectors embedded in the string are not widely used. And it would be a bit overkill to include a language detection heuristic (which would likely fail for short phrases). So should you let the user decide? Default to Japanese on a Japanese PC, otherwise leave it undefined?
- skhr0680 5y agoSet it by the device language, with a way to override that in your app’s settings
- lifthrasiir 5y agoThis is a good approximation, but it is still incorrect if, say, you are showing Japanese user names from a view for Chinese users.
- nikanj 5y agoAm I right in assuming fixing this in player names etc makes Chinese look wrong? It’s an easier problem if you know the whole page is Japanese, but how about things like game lobbies, where every username is in a different language?
- nitwit005 5y agoUltimately there are situations that aren't really fixable. People often mix languages in chat messages. There's some similar issues with right to left, and left to right text. You can give people a good default, and try to be smart, but some cases will always be ugly.
- yorwba 5y agoEither store the locale used when a user enters their name and then use it to mark up the text whenever you display the username, or simply use the system default, so Japanese users will see Chinese names with Japanese glyphs and Chinese users will see Japanese names with Chinese glyphs. Other users randomly get whatever.
- andrewl-hn 5y agoThe locale might be set to something completely different. A lot of programmers run their machines in English and not in their native language. One could use location to detect which variant to use, but that too wouldn’t work for, say, Chinese speakers in Japan. In ideal world we should use the locale of an input source (if the user sets their keyboard to Traditional Chinese we should use it for that fragment of text). However, operating systems and browsers don’t provide the input source locale API.
- rock_artist 5y agoSo to make sure I got it right: The issue is - when there's no proper glyph it uses a fallback? Or... The context of what glyph will be rendered depends on the document defined locale. (If it's the latter. it means it's impossible to quote another Asian language text within the same paragraph?)
- needle0 5y agoThe latter. In HTML you can specify a specific DOM element as being in a specific language so the browser can render it properly, but if the place you want to quote text isn't as allowing (eg. comment sections with no HTML allowed), there may be no way to ensure correct glyphs.
- lifthrasiir 5y agoIdeographic variation selectors plus a very large pan-CJK font may solve this issue in the future, but CJK fonts have already reached the OpenType limit of 65,535 glyphs so we are already running into technical issues.
- iforgotpassword 5y agoYou need to embed the quoted text in eg a span and add the proper lang tag.
- deleted 5y ago[deleted]
- tasogare 5y agoIt's a web browser only issue. In other cases such as a text processor document or local app a font is explicitly used by any run of text so there is no problem. This issue is the web being what it is, most of the time there is no font explicitly specified for text and the browser use a Chinese-looking font to display any Chinese characters.
- eska 5y agoIt’s not just a web browser issue. For example I’m transferring data in multiple Asian languages through some network API. I always need to specify the locale of the text data in a separate data field so that some UI program at the end can display the text correctly. And even then that’s not perfect, because that’s just the system locale instead of the IME locale.
- mjevans 5y agoThe initial three examples don't have a corresponding 'correct render' image next to them, so it's impossible for me to tell since they all render as the same character (which is incorrect given the lack of context). Checking the source, the page _is_ specifying language tags in the span, which I guess is supposed to help. My system just must not have fonts for those languages so I obviously can't even test them.
- needle0 5y agoThere may be issues displaying it on Firefox. Chrome and Safari seems to have displayed it correctly on my end. I'll find time to replace them with images so they appear correct regardless of environment. EDIT: Replaced with images.
- yorwba 5y agoFirefox will also display it correctly, but only if you have the fonts installed, the same as with other browsers.
- deleted 5y ago[deleted]
- simonlc 5y agoReally well done post and good idea with the title! - Simon from TGM :)
- iforgotpassword 5y agoMinor addition/clarification, just in case: > If the glyphs don’t exactly look like the Japanese result sample below, your code is displaying Japanese wrong. Maybe exactly isn't the right term here; it doesn't need to be pixel-perfect, there are still different font faces just like with western languages, for example one that's supposed to make them look more natural or hand written and one for print, etc. Also, afaict han unification was a mistake, but if you thought you only ever have 65535 code points available it might have been tempting.
- needle0 5y agoGood point. Will reword that part.
- wodenokoto 5y agoI think this page does a poor job of explaining that all 3 knives blade in the first example share the same code point in Unicode, but are to be displayed/rendered differently depending on which language it is shown as part of. It is there in the text, but it’s almost hidden between the lines. If I was a developer with no knowledge of Han characters or Han unification I would have to read two thirds of this article thinking I’m doing it right, so why am I reading this, e.g.: “but I am using the correct code point. It’s the character that the user entered!” or “I copy pasted it from a Japanese text, what do you mean I’m using the wrong character?” before reaching the “how to fix it” and even then I might not realize the root cause. With that in my mind I might not even make it to the part about how to fix the problem and learn that I am using the right character/code point, but it is still displayed wrong.
- needle0 5y agoI agree the page is somewhat roundabout in its current state since I went from the background to the symptom to the fix. Open to suggestions on rearranging the article so that more devs can implement fixes.
- wodenokoto 5y agoI would have the knives blades chart a little earlier, and make it very obvious that each character shares a code point (maybe have a code point column) and talk about in the text that yes, this is weird. 3 visually distinct Unicode characters share a single code point. For me this was very hard to wrap my head around the first time I encountered the problem. Maybe other people find it hard to understand in different ways. I believe that Unicode even claims that distinctly looking characters are to have their own code points, but similarly looking characters should share a code point (e.g, there is no French a and English a, even though they are pronounced differently. And Danish ø and Swedish ö are pretty much the same pronunciation but differently written, so they don’t share a code point.)
- needle0 5y agoThanks. I reworded it a bit and put more emphasis on code points.
- wodenokoto 5y agoWonder how well HN and my phone handles this. There are supposed to be Unicode code points that indicate which locale a character is supposed to be displayed in. If things are well thought out, my phone should add them automatically and HN should keep them and your browser should render it correctly On an iPhone using, Chinese simplified keyboard: 刃 Japanese keyboard: 刃 So that didn’t go very well. When choosing the character on my Chinese keyboard it is displayed with correct Chinese strokes but turns into the Japanese version in the text box. I’m guessing for most of you reading, both will appear Chinese. EDIT: Someone better than me at wrangling unicode can maybe try out the variation selectors, and print the correct variations in a comment. I think it would have been neat if my keyboard ime did it for me :) https://en.wikipedia.org/wiki/Variation_Selectors_(Unicode_block) https://en.wikipedia.org/wiki/Variation_Selectors_(Unicode_b...
- aikinai 5y agoThey both appear Japanese for me on mobile Safari. But Japanese is my second preferred language on the device (after English), so that’s probably why.
- wasmitnetzen 5y agoYep, Simplified Chinese for me (with LANG=en_US.UTF-8).
- skneko 5y agoBoth appear Japanese for me. Locale = es_ES.UTF-8
- thrdbndndn 5y agoI'm a native-CJK user myself and well aware this phenomenon, but honestly it's not really that bad in most of cases due to the following reasons: 1. Websites that are in Japanese are likely tagged with lang=ja already. So they will display fine. Unfortunately, this practice seems to be less followed by Chinese sites. I checked a few top sites, qq.com do have lang=zh-cn, while baidu.com and sina.com.cn don't. 2. Majority of UI elements in OS will prioritize the display language you set when choosing variants. This means, if the users are reading content in Japanese while also using Japanese UI, the glyphs would be correct. Of course, this will cause problem if a Japanese is reading Chinese or vice versa, but such scenario is in minority. Another scenario, which I think is more common, is when someone is using a Latin-language UI. For example, lots of my (Chinese/Japanese) friends are using English UI while reading Chinese/Japanese a lot. The OS in this case will default to one variant (I believe Apple by default would choose Japanese) and therefore display another language's glyphs wrong (side note: for web pages, desktop browsers often have their own font/glyph fallback logic above the OS one). 3. Most of people are just not sensitive to such thing. I pointed it out to lots of people (when due to their setting, some glyphs are displayed wrong, like 门), and they can't care less. Also, there is no simple "fix" if you have multi-language content. Without manually assign <lang> tag to every single string, you can't display both Japanese and Chinese correct at the same time. It isn't worth the hassle for just a few phrases in text. A good example is Wikipedia, they have templates for all kinds of languages so you can display them correctly even if it's just one Japanese word on, say, English Wikipedia. And Wiki editors do use them all the time!
- eloisant 5y agoWhat you're saying is true (although I'm not sure about 3, all the Japanese people I've talked to are annoyed by that), but it really sucks that we're dealing with problems that were supposed to be fixed by Unicode. Han unification have been a huge mistake, to save a few thousands of characters, and now we keep piling on more and more stupid emoji.
- dotancohen 5y ago> Han unification have been a huge mistake, to save a few thousands > of characters, and now we keep piling on more and more stupid emoji. This is my exact problem with Unicode. I've very grateful for the efforts that they have made in the past, but the change from "spare valuable codepoints at the expense of causing ambiguity in text" to "assign a new codepoint to every cartoon permutation of intangible nouns" is infuriating.
- kiryin 5y agoAs a Japanese learner, this has been a massive disappointment in unicode for me, and a pain in my ass. It has sort of formed into a challenge for me, trying to get the characters to display consistently on all of my devices. Believe it or not, even with pango configured to always show the japanese variants, and fontconfig set to always prefer the JP font, some applications like Firefox find a way to mess it up. Can't blame them much though, han unification is a huge mess and designed by someone who I can only posit to be entirely brainless. There aren't many characters that are affected, you aren't even saving any considerable amount of codepoints. It's just west-centricism and lack of knowledge on the subject.
- adrian_b 5y agoThe Han unification was done because at that time they hoped that the size of Unicode characters will be limited to 16 bits. Separate sets of Han characters cannot be encoded in the 16-bit space, but they could have been easily encoded in the current 32-bit space. Nevertheless, I have never found this to be a problem in practice, because I have always taken care to have good separate typefaces for Japanese, Traditional Chinese and Simplified Chinese. In documents that I create or modify, I apply styles with the appropriate typeface. The only possible problems are with Web pages, but the good browsers allow you to configure typefaces for each language and I always configure the correct typefaces. If the Web page does not specify correctly the language, it might be displayed wrongly, but this is only one of the many stupid things that can be done by a Web page designer that can make that page look ugly when rendered on other computers.
- aikinai 5y agoThe fact that you have to take such care is exactly the problem.
- adrian_b 5y agoI agree that the fact that I must not forget to configure typefaces per language whenever I install a new browser, while for Chrome you must also install the "Advanced Font Settings" extension before it even becomes possible to choose e.g. a Japanese font, is annoying. To avoid such configuration work when you prefer better looking typefaces instead of some standard system defaults would require a standardization of how to notify the applications about the association between certain typefaces and languages, e.g. by some environment variables or by some standard locations for the font files, depending on language.
- suction 5y agoDoes this also explain why alphabetic text in Japanese apps and websites often looks so horrible? Like very wide characters with way too much space in between them?
- tasogare 5y agoNo, the wide characters (it's their name) are special code points. I guess it exists because someone wanted to be able to use one letter in place of a Japanese character while using the same width. The "normal" letters are called "half-width" here.
- lifthrasiir 5y agoFull-width characters are relics from multiple legacy character sets. For example JIS X 0208, the primary Japanese two-byte character set, has a set of alphanumeric characters in the row 0x23, but their widths are not specified and it is totally possible to map them into half-width characters when no other character sets are in use. However it is most commonly paired with JIS X 0201 which is a single-byte character set with their own alphanumeric characters, so anything from JIS X 0201 is made half width and anything from JIS X 0208 is made full width to simplify implementations. This practice got stuck and subsequently followed by Unicode. Same for other languages.
- jhanschoo 5y agoAlphabetic text in Japanese fonts are primarily designed for documents mainly in Japanese with the occasional Latin script jargon. There's a variant (full-width) that's sometimes used that indeed is very wide, made to be of the width of the Japanese kanji, but even the proportional ones are pretty light and widely spaced (which results in better typography in mainly-Japanese documents)
- needle0 5y agoNewspaper websites may also have years-old internal typesetting rules, carried over from paper, that mandate alphabetical text must appear in full-width (double wide). They look ugly even to native Japanese, and some newspapers have gradually been learning to break out of it.
- lovasoa 5y agoWasn't the whole point of Unicode to have a single encoding that could represent all languages unambiguously so that you don't need any meta-information to display a string ? Is there a reason why they chose to represent characters that are obviously different with the same code point ? Everyone would find it outrageous if they decided to have a single character for the russian м and the english m just because they have the same greek origin...
- Sniffnoy 5y agoThe reason why is that Unicode was originally 16-bit, and there was no way they could fit everything into 16 bits without CJK unification. Of course later it turned out there was no way they could fit everything into 16 bits anyway, and so they were forced to expand it, and so we now both have a larger Unicode (with all the messes that's caused) but also still have CJK unification...
- lovasoa 5y agoIs it too late to add separate code points for the chinese, japanese, and korean versions of the han characters ?
- lifthrasiir 5y agoIn principle there is a designated "disunificiation" procedure when it's desirable. More accurately speaking, each CJK character is thought to represent not a single or a few glyphs listed in the code chart but rather a glyphic subset, and the disunification splits that set into partitions. But this is generally applied to a few selected characters and only when it's safe to do so. Massive disunification was to my knowledge never suggested or proposed, and that would surely prompt a large scale disruption throughout CJK users (say, how about existing texts?).
- rrobukef 5y agoSo it's possible to modify the skin tone of emoji but impossible to disunify CJK characters? There are RTL modifiers for Arabic languages, it's impossible for CJK? It shouldn't be harder than existing unicode handling.
- lifthrasiir 5y agoUnfortunately this (and linked) article only represents Japanese issues. If you blindly apply these suggestions Chinese or Korean users may have issues. I'll list Korean issues below primarily because I'm Korean, but you may want to interview actual CJK users (one of each, not a single user) for testing. > Line breaking rules This should link to W3C Requirements for CJK Text Layout [1]. The Wikipedia article alone doesn't fully describe the complexity of CJK typography. CJK languages are common in that they all have classes of punctuations that can't be separated by a newline. But there is one more thing to consider for Korean: both word-based breaking and character-based breaking is possible depending on the context. The general rule is to use word-based breaking for larger texts and character-based breaking for smaller texts, but there is no clear threshold so you really want to consult Korean users for testing. [1] https://www.w3.org/TR/clreq/ https://www.w3.org/TR/clreq/ (Chinese), https://www.w3.org/TR/jlreq/ https://www.w3.org/TR/jlreq/ (Japanese), https://www.w3.org/TR/klreq/ https://www.w3.org/TR/klreq/ (Korean) > Messaging Apps: Do not directly hook to the Enter key to submit messages This advice is also problematic. In pretty much all Japanese and most Chinese IMEs they should go through candidate windows so pressing Enter should not submit messages, but in some Chinese and virtually all Korean IMEs there is no automatic candidate window and pressing Enter should submit messages. In the ideal world detecting a newline as suggested by the article should have solved this issue, but that got complicated by clueless pan-CJK IME implementations. They generally assume candidate windows even for Korean, so they do not commit texts on Enter and that's very inconvenient for Korean users. Therefore it is rather recommended to detect a newline by default, but also have an option to submit messages on Enter.
- needle0 5y agoI updated both sections according to your suggestions. Thanks!
- needle0 5y agoWas notified from someone else about the isComposing attribute -- https://developer.mozilla.org/en-US/docs/Web/API/KeyboardEvent/isComposing https://developer.mozilla.org/en-US/docs/Web/API/KeyboardEve... At least for web stuff, do you think checking for this before treating the Enter key as Submit would work in both IMEs with and without input buffers?
- brigandish 5y ago> Japanese text written in incorrect glyph sets will stand out similarly to any native speaker of Japanese, and will give off a connotation that whoever developed this app does not care about this (often large) subset of the global user population. More likely they'll think the content was written by a non-native Japanese speaker, judge whether that makes you trustworthy or not (based on personal experience or stereotypes or prejudice, probably a bit of all three (we're all human)) and then not buy from you. A good example would be Amazon listings in Japanese that Japanese people can tell were almost certainly written by someone Chinese, and then decide not to buy. If you want the cash, get a proper translation. Ironically, Japan is filled to the brim with incredibly poor English and abounds with stories of native English speakers' translations and corrections being disregarded because "it doesn't sound right"… to someone who can't string a legible English sentence together.
- xvilka 5y agoIt's more of a problem in pure text apps rather than the Web. For example, in editors (not the rich text ones), console, interface elements. But yes, it is a problem for people who knows (or learns) and uses multiple languages at once, e.g. English, Chinese, and Japanese.
- hannob 5y agoI didn't know about this, but can't help to think this sounds like a bug in unicode to me. If these characters are different then why does unicode assign one codepoint to them? Wasn't the promise of unicode to exactly not do this kind of thing? Can this be fixed? New character code for ambiguous characters could be assigned, of course this would require manual conversion (with knowledge of the variant) for existing data, but at least it would make this issue go away moving forward (and unconverted legacy data would be "just as bad" as it used to be, so no loss).
- oleganza 5y agoThis issue with Han Unification was a big reason for stalled adoption of Unicode/UTF-8 in Ruby for years. UTF-8 by default came to Ruby after 1.9 where they've added thorough support for variety of encodings, so that UTF-8 is not the only option.
- oleganza 5y agoHan Unification started in the 90s when computers were big, memory small, UTF-8 did not exist and people were trying to fit all characters in a reasonable amount of codepoints. Today with variable-length encoding of UTF-8 and video streams over 5G, supporting all variants as distinct codepoints, and patching text search and sorting with more "normalization" algorithms would not be a problem at all.
- afiori 5y agoin my opinion utf8 should have been a bigger variable length encoding, today it is: 0xxxxxxx 110xxxxx 10xxxxxx 1110xxxx 10xxxxxx 10xxxxxx and 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx the only reason not to push those last bits and add 111110xx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 1111110x 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 11111110 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx and maybe even 11111111 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx is utf-32, they should have dropped it and solve the codepoint problem this way.
- lifthrasiir 5y ago
- BoppreH 5y agoThat's extremely interesting, if not depressing. So, if I have to display user-entered text (usernames, posts, comments, messages, form data, etc), and I want to do The Right Thing™: - I cannot rely on user locale, because it might be set to something generic like English, or the user may be bi-lingual. - I cannot rely on location, because the user may be traveling to a different CJK region, or somewhere else altogether. - I cannot set a single lang: attribute for the whole page because it'll be wrong for the other two languages. - The string alone is not sufficient to identify the language because you can write valid sentences in different CJK languages with the same codepoints. - I cannot have a per-user language setting, because users may be bi-lingual. What does that leave me? A dropdown list "C/J/K/Other" besides every single text field? I'm chucking this on my pile of examples of software development being hopelessly broken by design, along with "unix time is non-monotonic and discontinuous at random" (hint: what's the unix time exactly 1e8 seconds, ~3 years, from now? Answer: it's up to the astronomers[1]!). [1]: https://en.wikipedia.org/wiki/Unix_time#Leap_seconds https://en.wikipedia.org/wiki/Unix_time#Leap_seconds Edit: actually, even the dropdown list is insufficient because it only allows one language per string! How is a Japanese user asking for help learning Chinese supposed to write?
- zokier 5y ago> unix time is non-monotonic and discontinuous at random Well, depends on your definition of unix time; if you use time() as the definition then it is actually monotonic because the integral part only repeats on leap seconds?
- dathinab 5y agoThe problem is btw. not specific to Han unification/asian languages. It's that for every western language if you use screen readers.
- the_other 5y agoMy thought whilst reading the article was that the Han unification would actually help screen readers. IIUC The meaning of the glyph is the same across all the languages, so the screen reader will get the correct meaning and can present it according to local settings. The problem with the European languages is that the different characters (letter variations, accent variations) can change the meaning of the word they're part of. Or have I misunderstood?
- squaresmile 5y agoThis can also be a problem in chat app. I used en-US on windows (which defaulted to the zh variant) and someone else used ja-JP and I was wondering why the character was different. Took a while to notice that we were seeing two different things on our screens. We also have a website about a Japanese game using a Japanese font except for 0x9bd6. The font's 0x9bd6 is the CN variant and its 0xe001 is the JP variant of 0x9bd6. Fun times. Like others said, on the web, you pretty much have to manually assign lang to every single thing. We just added support for CN/TW/KR text. I should come back and check 0x9bd6 in the other versions ...
- euske 5y agoThis is the reason why Adobe PDF isn't relying on Unicode. Adobe products has a huge presence in Japan since 90s and they had to appeal to the printing industry, which is very anal to this kind of issues. So they ended up using a separate encoding for every language. Today, CJK letters in PDF are encoded in Adobe-GB1 (mainland China), Adobe-CNS1 (Hong Kong), Adobe-Japan1 and Adobe-Korea1 respectively. Not the cleanest way, but it gets the job done.
- lifthrasiir 5y agoNote that they are now adopted by the Unicode Ideographic Variation Database [1] among other variation databases. [1] https://unicode.org/ivd/ https://unicode.org/ivd/
- makeitdouble 5y agoThanks for the pointer, that's pretty interesting. Looking at their doc [0] it seems they used their Adobe-Japan1 to wrap a much more wider set of characters than any single encoding standard, including ligatures, vintage encodings etc. It seems to be a pretty big work and kinda fits with the image of PDF handling being such a monumental beast. [0] https://github.com/adobe-type-tools/Adobe-Japan1/ https://github.com/adobe-type-tools/Adobe-Japan1/
- ksec 5y agoAdobe gets lots of stick for its subscription and malware like Creative Cloud. But they do spend huge amount of resources on CJK fonts, layout and encoding. And part of the reason why I like PDF. ( Behind a Paywall ) https://ken-lunde.medium.com/my-28-years-of-adobelife-e97e703fd924 https://ken-lunde.medium.com/my-28-years-of-adobelife-e97e70...
- powerapple 5y agoFor Chinese, the font can change the writing slightly. For example, 刃(blade) can be any of these (Japanese, Simplified Chinese and Traditional Chinese). Actually I would consider the Japanese version in the article the traditional Chinese version: https://duckduckgo.com/?q=%E5%88%83+%E4%B9%A6%E6%B3%95&iax=images&ia=images https://duckduckgo.com/?q=%E5%88%83+%E4%B9%A6%E6%B3%95&iax=i... At least for Chinese, the difference is font, they are all valid writing for the character. Different writing style can cause these minor difference as well: http://qiyuan.chaziwang.com/pic/ziyuanimg/E58883.png http://qiyuan.chaziwang.com/pic/ziyuanimg/E58883.png If you look at right side of above image, you can tell how the same character is written in different writing style It is less a problem for Chinese. Our brain has trained to read them, I would recognize the Japanese version, Simplified Chinese version, and Traditional Chinese version without noticing the difference. But I can imagine it can be a problem for Japanese, and other people do not read Simplified Chinese. Having the locale explicitly set to a country and load the correct font make a lot sense here.
- peacefulhat 5y agoI'm grateful for han unification because I can search Chinese words I only know in Japanese.
- ksec 5y agoI am surprised how many comments here never heard of Han Unification. The problem is not new, and some of us have been ranting about it for more than a decade. From the UTF-8-Everywhere Manifesto in 2012 on HN [1], And a search [2] on HN dates back to 2010. I am also surprised at the support this problem now has. At least on this thread. Generally speaking Han Unification problem dont get much if any support on HN. Not even empathy. In the name of having Unicode becomes king they would much rather sacrifice the CJK language. The answer or replies were always, it is "glyph" problem, not "code" problem. Stop asking Unicode to solve it. patio11 aka Patrick McKenzie from Stripe has been the most vocal critics of Han Unification. Sums it up far better than I could, quote [3]: >Reason the Han unification debate in Unicode got so acrimonious, and why lots of Japanese people carry a chip on their shoulder about it to this day. >"Sorry, grandma, I know you've been sort of attached to your name for the last 80 years, but the white folks find it inconvenient for their computer systems. Don't worry, they promise they'll make something close for you." >Many of the clients of my ex-day job are married to legacy encodings like Shift-JIS precisely because they do think that their customers and students have a "right" to having their names written correctly. As mentioned in my other reply, Adobe gets lots of stick for its subscription and malware like Creative Cloud. But they do [4] spend huge amount of resources on CJK fonts, layout and encoding ( They have their own separate Encoding for each CJK language instead of using Unicode ). Part of the reason why I like PDF. [1] https://news.ycombinator.com/item?id=3906253 https://news.ycombinator.com/item?id=3906253 [2] https://hn.algolia.com/?dateRange=all&page=8&prefix=false&query=Han%20Unification&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=8&prefix=false&qu... [3] https://news.ycombinator.com/item?id=1438749 https://news.ycombinator.com/item?id=1438749 [4] https://ken-lunde.medium.com/my-28-years-of-adobelife-e97e703fd924 https://ken-lunde.medium.com/my-28-years-of-adobelife-e97e70...
- flubert 5y ago>"Sorry, grandma, I know you've been sort of attached to your name for the last 80 years, but the white folks find it inconvenient for their computer systems. Don't worry, they promise they'll make something close for you." Is there a resource to read more about this? I don't get that vibe from things like: https://www.unicode.org/versions/Unicode3.0.0/appA.pdf https://www.unicode.org/versions/Unicode3.0.0/appA.pdf
- YeGoblynQueenne 5y ago>> However, this issue is much more than the difference between, say, the lowercase A with the overhang (a) or without (α). Yes but actuallly "α" is the Greek character alpha, whreas "a" is the Latin character "a". So if you displayed "α" as "a" to a Greek person that, too, would look αλλ ωρονγ.
- zokier 5y agoThis is a tangent, but has there been any ideas around making a stroke-based encoding of Japanese/Chinese writing systems?