10 ms·
As the story mentions regarding the off symbol (a circle), there are many visually identical code points that have different semantic meanings. But in this case
by hackuser 10y ago
As the story mentions regarding the off symbol (a circle), there are many visually identical code points that have different semantic meanings. But in this case, they added an additional semantic meaning to an existing code point.
So which is it? Does each code point represent a visual image? A semantic meaning? Both? It depends? Something else?
I've tried to decipher that on my own and only learned that the answer to these sorts of questions are complicated, because it's very complicated to represent all written human language via one set of rules.
So I know some of the answers to my questions above, but I'm hoping someone with real expertise can provide the fundamental rules/policies - if there are any.
- mrb 10y ago"So which is it? Does each code point represent a visual image? A semantic meaning? Both? It depends? Something else?" Well the answer is clear: each code point represents one visual image, to which is associated one or more meanings.
- coddingtonbear 10y agoThat is not the compromise struck, though; there are even many Cyrillic glyphs that are visually identical to those in Latin, but assigned differing codepoints.
- jhanschoo 10y agoThere are multiple reasons for that, one of which is compatibility with previous encodings and standards. If a previous encoding Unicode wanted to be compatible with encoded these as different characters, Unicode needs these to have separate code points for them too.
- hvidgaard 10y agoThat is the surefire way to incorporate complexities from 2 different systems into 1.
- xyproto 10y agoBeing able to easily check if a letter is between 'a' and 'z' in code is an advantage. This is only possible if the codepoints are sequential.
- hvidgaard 10y agoI didn't dispute that. I just state that trying to remain compatible for the sake of being compatible is a great way to design a convuluted and difficult to understand standard.
- gpderetta 10y agoOf course, but lack backward compatibility is a great way to make sure a standard is not adopted. For example he reason that UTF-8 'won' is that it has a great backward compatibility story with other ASCII based encodings and systems.
- bambax 10y agoIs it? Couldn't Unicode have pointers or links, where a codepoint "exists" with no content and only links to another? (I don't know anything about Unicode, so maybe it already has that.)
- edent 10y agoSemantically, yes. In the code tables you'll see that that "opposite" symbols have links to each other. Programatically, it is much easier to say "does a character lie between 0x12 and 0xBC" than to create a function like `isSymbolForTrafficInEurope()`
- Natanael_L 10y agoPerhaps Unicode could just have tables listing all relevant sequences of symbols, instead. So "latin letters lowercase" would list the codepoints for a-z in order, for example. Would no longer matter if the codepoints themselves are sequential or not. (And relevant to my country, "Swedish characters lowercase" would map to latin letters lowercase + åäö.)
- deleted 10y ago[deleted]
- iopq 10y agoYes, but they have alternate italic forms, for example. Sure, some one of the glyphs like с doesn't have an alternate italic form. Since the other ones do, it would be weird to only assign a separate codepoint to some of them and overlap the others. It would be a workable solution, but still weird.
- wodenokoto 10y agoTry looking up han-unification and its justification and you'll see the exact opposite approach to encoding characters into unicode. For CJK characters, they unified all semantically similar han-characters, even when they have visual forms that are quite different between Japanese, Chinese and Korean. If you want to write Japanese and Chinese in the same document, you need to mark up the section to tell the system that renders it, to render different visual forms for similar codepoints depending on whether they are used in Japanese or Chinese.
- thaumasiotes 10y ago> For CJK characters, they unified all semantically similar han-characters, even when they have visual forms that are quite different between Japanese, Chinese and Korean. This isn't true. 青 and 靑 are the same character written differently; they have their own codepoints. Ditto for a huge number of simplified Chinese characters; 语 is mainland Chinese and 語 is the same character in Japanese.
- wodenokoto 10y agoIt is true for lots of characters (so I guess I was being a little hyperbolic when I said "all"), and you cannot rely on choosing the correct code points in order to have a text display Japanese or Chinese. You need to tell your rendering program (often through choice of font) if things are to be rendered with Japanese or Chinese forms. I wouldn't know how to show you examples here, as 直 will 直 display the same since they have the same code point, but different number of strokes in japabese and chinese. https://en.m.wikipedia.org/wiki/Han_unification https://en.m.wikipedia.org/wiki/Han_unification
- rspeer 10y agoAren't they putting the disunified characters into the U+2xxxx plane now? Han unification is generally seen as a bad choice in retrospect, but it was something Unicode had to do when it looked like 2^16 codepoints were all they were going to get.
- EdiX 10y ago> So which is it? Does each code point represent a visual image? Look it's pretty simple, every code point represents a semantic meaning, except for: 1. those characters who also encode the width of their visual image (U+FF00..FFEF) 2. the one that means 'unknown' (U+FFFD) 3. those characters that change their visual representation depending on their position in the word (U+FB50..U+FDFF,U+FE70..U+FEFF) 4. those that change the visual image of another code point (U+FE00..U+FE0F) 5. those characters that have a visual image as their semantic meaning (too many to list) 6. those that are designated to have no semantic meaning at all (U+FDD0..U+FDEF) 7. those that have a meaning only in pairs (U+D800..U+DFFF) 8. miscellaneous
- tragomaskhalos 10y agoWhy does this list remind me of https://en.wikipedia.org/wiki/Celestial_Emporium_of_Benevolent_Knowledge https://en.wikipedia.org/wiki/Celestial_Emporium_of_Benevole... ? :)
- EdiX 10y agoBecause Unicode is the Celestial Consortium of Benevolent Encoding.
- PeterisP 10y agoWhy would #1 #3 #4 and #5 not apply? "every code point represents a semantic meaning" is completely consistent with the notion that some code points e.g. have differing visual representation depending on their position in the word.
- coldtea 10y ago>As the story mentions regarding the off symbol (a circle), there are many visually identical code points that have different semantic meanings. So? How is that different from any regular character in real life? 101 for example means the number 101, an introductory class in university, slang for "anything introductory" in general, etc. And let's not get started on the meanings of letters, e.g. a and e.
- jomamaxx 10y ago"many visually identical code points" The emoji code points can be represented differently on different systems given their meaning. So it makes sense to have different emojis for different 'meanings'. The 'moon' switch here does no mean 'moon' - it means 'standby' or whatever. It may look noticeably different on different systems. Think from a design perspective: you have 5 emojis to represent 'clouds, sky, earth' etc. - and the a different set of 5 to represent 'on, off, sleep, shutdown'. Those icons will be markedly different in terms of representation, groupings, colour coding, underlying functionality if they are integrated into an experience in any meaningful way. Text your car with the 'shutdown' symbol to tell it to shut down. Your bot texts your friend with a moon symbol to tell him you're asleep. Or whatever.
- xg15 10y agoBut that doesn't explain the inconsistency in the current case. So if a system wants to render "on" differently than "straight vertical line", that's possible. However, if "off" should be rendered differently than "circle", that's not possible. (Or only possible with out-of-band information or modifier characters which would still have to be defined)
- msbarnett 10y agoYeah. It's a mess. If you want to write a document in Japanese that talks about a Chinese character which is written differently than its Japanese version, you can, or can't, achieve this in Unicode, depending on the character, its history, and the mood of the consortium the day it was assigned. The reality is that Unicode is governed by people, some of those people are grumpy reductionists who push for a minimum of symbols and a maximum of meaning-overloads, and others are more liberal and tend to advocate the opposite, and the result is a compromise, and is in areas very messy.
- PeterisP 10y agoDo note that they did not include a generic "off" symbol, they included the IEEE 1621 off symbol - which must be rendered as a circle; while on the other hand the IEEE 1621 on symbol must be rendered in a manner that is often different from just "straight vertical line" in particular regarding the corners of that line.