4 ms·
I'm really curious about how this will be implemented from a software standpoint. Will the language use the same ISO (or whatever body determines them) browser
by swrobel 8y ago
I'm really curious about how this will be implemented from a software standpoint. Will the language use the same ISO (or whatever body determines them) browser language code like en-US or will there need to be two separate codes that exist in parallel, one for the old cyrillic version, and one for the new latin one.
Are there any modern examples of this sort of transition having to be implemented? I think we all think of the set of possible language codes as something that has been static for a long time, but this shows that it can be rather fluid.
- rspeer 8y agoBCP 47 language codes [1] have three major parts: the language, the script, and the region. Canonically, the language is two or three lowercase letters, the region is two capital letters or three numbers, and the script is four letters in title-case. These values can be filled in from context. "en-US" means the same thing as "en-Latn-US" because it's always written in the Latin alphabet. "ja" (Japanese) means the same thing as "ja-Jpan-JP" (Japanese, as used in Japan, written in the combination of scripts that is unique to Japanese). You could refer to romanized Japanese as "ja-Latn", though this is really rare. But there are a few language codes where you should specify the script. Particularly Serbian, where the Latin and Cyrillic scripts coexist. In that case, you distinguish them as "sr-Latn" and "sr-Cyrl". So the language codes for Kazakh will be "kk-Latn" and "kk-Cyrl". Which one "kk" means by default may change at some point. [1] https://tools.ietf.org/html/bcp47 https://tools.ietf.org/html/bcp47 - if you find yourself referring to "ISO language codes" in the present day, this is what you actually mean.
- stordoff 8y ago> "ja" (Japanese) means the same thing as "ja-Jpan-JP" (Japanese, as used in Japan, written in the combination of scripts that is unique to Japanese). Incidentally, I noticed that Facebook has both ja-JP and ja-KS (for Kansai-ben) recently, which I haven't seen elsewhere before.
- rspeer 8y agoLanguage codes made up off the top of someone's head are my peeve. If anyone else wanted to support Kansai dialect, they wouldn't necessarily (and probably shouldn't) make the same decision to pretend that it's a country with code KS. It's not like the standards left them with anything to work with, but "ja-x-kansai" would have been quite acceptable. At least it's not as bad as code I've seen that used "zh-SC" and "zh-TC" to represent Simplified vs. Traditional Chinese, because they either didn't know about "zh-Hans" vs. "zh-Hant", or didn't leave room for script codes in their database. (In my post I neglected to even mention Chinese, by far the largest example of a two-script language.) If you read the codes "zh-SC" and "zh-TC" literally, they're distinguishing whether it's "Chinese as used in Seychelles" or "Chinese as used in the Turks and Caicos Islands". And it's not as bad as the OpenSubtitles language code "ze", which after some examination, I have to conclude means "this might be Chinese or might be English, we're not sure, we found it on a shoddily pirated DVD".
- jwilk 8y agoEven if they didn't have room for script codes, they could still use zh-CN / zh-TW.