6 ms·
Wait, "ত + ্ + = ৎ" is nothing like "\ + / + \ + / = W". The Bengali script is (mostly) an abugida. Ie, consonants have an inherent vowel (/ɔ/ in the case o
by Epenthesis 12y ago
Wait, "ত + ্ + = ৎ" is nothing like "\ + / + \ + / = W".
The Bengali script is (mostly) an abugida. Ie, consonants have an inherent vowel (/ɔ/ in the case of Bengali), which can be overriden with a diacritic representing a different vowel. To write /t/ in Bengali, you combine the character for /tɔ/, "ত", with the "vowel silencing diacritic" to remove the /ɔ/, " ্". As it happens, for "ত", the addition of the diacritic changes the shape considerably more than it usually does, but it's a perfectly legitimate to suggest the resulting character is still a composition of the other two (for a more typical composition, see "ঢ "+" ্" = "ঢ্").
As it happens, the same character ("ৎ") is also used for /tjɔ/ as in the "tya" of "aditya". Which suggests having a dedicated code point for the character could make sense. But Unicode isn't being completely nutso here.
- elFarto 12y agoI don't understand Bengali at all, but the character ৎ does have it's own Unicode codepoint (U+09CE, BENGALI LETTER KHANDA TA). It was introduced in Unicode 4.1 in 2005.
- chimeracoder 12y agoYes, this is what the article says: > Until 2005, Unicode did not have one of the characters in the Bengali word for “suddenly”. The codepoint does exist now, but it took ten years before it was included.
- Manishearth 12y ago... and emoji was added to the Unicode standard in 2010. The title of your article is fallacious :/ (though I agree with most of it)
- ubasu 12y agoNative Bengali here. The ligature used for the last letter of "haTaat" (suddenly) is not the same as the last ligature in "aditya" - the latter doesn't have the circle at the top. More generally, using the vowel silencing diacritic (hasanta) along with a separate ligature for the vowel ending - while theoretically correct - does not work because no one writes that way! Not using the proper ligatures makes the test essentially unreadable.
- eridius 12y agoI don't understand Bengali at all. I'm trying to understand your second sentence though. When you say "no one writes that way", do you mean nobody hits the keys for letter, followed by vowel-silencing diacritic, followed by another vowel? Or do you mean the glyph that results from that combination of keystrokes doesn't match how a Bengali speaker would write it on paper? If it's the latter, isn't that an issue for the text input system to deal with? Unicode does not need to represent how users input text, it merely needs to represent the content of the text. For example, in OS X, if I press option-e for U+0301 COMBINING ACUTE ACCENT, and then type "e", the resulting text is not U+0301 U+0065. It's actually U+00E9 (LATIN SMALL LETTER E WITH ACUTE), which can be decomposed into U+0065 U+0301. And in both forms (NFC and NFD), the unicode codepoint sequence does not match the order that I pressed the keys on the keyboard. So given that, shouldn't this issue be solved for Bengali at the text input level? If it makes sense to have a dedicated keystroke for some ligature, that can be done. Or if it makes sense to have a single keystroke that adds both the vowel-silencing diacritic + vowel ending, that can be done as well. --- If the previous assumption was wrong and the issue here is that the rendered text doesn't match how the user would have written it on paper, then that's a different issue. But (again, without knowing anything about Bengali so I'm running on a lot of assumptions here) is that still Unicode's fault? Or is it the fault of the font in question for not having a ligature defined that produces the correct glyph for that sequence of codepoints?
- ubasu 12y agoIt has to do with how the text is rendered. For example, if you see the Bengali text on page 3 of this PDF: http://www.unicode.org/L2/L2004/04252-khanda-ta-review.pdf http://www.unicode.org/L2/L2004/04252-khanda-ta-review.pdf it is unreadable and incorrect Bengali. ;-)
- Tomte 12y agoBut Unicode emphatically does not define a rendering, a glyph. To me it sounds like the rendering needs to be fixed, not Unicode.
- 12y ago
- NelsonMinar 12y agoI'm a bit confused too. There's a principled argument in Unicode for when a glyph gets its own codepoint vs when it's considered sufficient to use a combining form. I don't know Bengali at all so can't comment on this case, although given the character now is in Unicode I guess the argument changed over time. Somewhere buried in the Unicode Consortium notes is an explicit case for the inclusion / exclusion of this character, it'd be interesting to find it.