3 ms·
What's wrong with just wrapping UTF-8 text in lang spans? Browsers and Electron will just do the right thing, no?
by amanwithnoplan 4y ago
What's wrong with just wrapping UTF-8 text in lang spans? Browsers and Electron will just do the right thing, no?
- lmm 4y agoEven assuming that works, getting it to make its whole way through your whole tech stack is no easier (and more HTML-specific) than having spans with their own encoding.
- astrange 4y agoSince when can HTML pages have multiple encodings at the same time?
- TheDong 4y agoThis isn't about multiple encodings, this is about the one encoding, UTF-8, being able to represent multiple languages. Which is mostly can, except for han characters. In my reply here, I can type in english and russian at once. Привет, мир. Yet, if I try to type chinese on one line, and japanese on the next, I cannot do it. Hacker news does not let me enter "lang" tags, so I can only type either the chinese or the japanese variant of a kanji.
- astrange 4y agoParent post says > spans with their own encoding. so yes it is about pages with multiple encodings. A span's smaller than a page! The Unicode answer is "variation selectors" which are used for some historical variant kanji, but not for whole language switching. I suppose they could be used for that too though.
- lmm 4y agoI don't know about HTML specifically; I meant spans as a general concept rather than a literal <span> tag. If the web stack has actually implemented mixing languages on the same page to the point where you can use it in a "normal" application then that's very cool (and if they've done it with their lang tags rather than by allowing mixed encodings, well, fine), but I've yet to see a site that actually has that up and running.
- Symbiote 4y agoI can't read it, but there will be hundreds of articles on Japanese Wikipedia covering Chinese literature etc that have text in both languages, all in Unicode.
- lmm 4y agoHad a look, they have a macro for it. Cool!
- Symbiote 4y agoAlso, <html lang="ja"> Japanese text ... <span lang="zh"> Chinese quote </span> ... </html> is much easier than mixing encodings. With the above entirely in Unicode, it will be handled reliably by anything that can handle Unicode, and is still reasonably readable even if the Chinese text is shown in a Japanese font. Reading just the fourth line without the third will still show something 'OK'. Mixing Shift_JIS and Big5 sounds like a recipe for corruption, but something similar was done in an old Russian and Japanese encoding: https://en.wikipedia.org/wiki/Shift_Out_and_Shift_In_characters https://en.wikipedia.org/wiki/Shift_Out_and_Shift_In_charact... RFC 2482 was a Unicode adaptation of this, but it was deprecated 12 years ago: https://www.rfc-editor.org/rfc/rfc6082.html https://www.rfc-editor.org/rfc/rfc6082.html