4 ms·
Oddly, the following more modest proposal hasn't gotten much traction: characters that share history but have divergent graphical representations in the various
by yuubi 12y ago
Oddly, the following more modest proposal hasn't gotten much traction: characters that share history but have divergent graphical representations in the various dialects of alphabetic script shall share codepoints, and a mechanism beyond the scope of Unicode (like lang attributes or plain guesswork) shall be used to decide whether a given codepoint means L or ᴫ or Λ or whatever.
- fenomas 12y agoActually I think that's roughly how things work today. It's not my area, but here's how this was explained to me recently by a colleague (the lead font guy at Adobe Japan): There are two ways of dealing with glyphs that share code points. The first is TTC (truetype collection) fonts. A TTC is basically one set of glyphs with several sets of mappings (i.e. which code point maps to which glyph). When you install it, assuming your computer groks ttc, your system shows you a separate font for each mapping. Taking for example Source Han Sans, which adobe just released - if you go to the download page[0] and get the complete version (the "OTC" one), you get a bunch of files like "SourceHanSans-Bold.ttc". If you install one of them you'll see four new fonts: "Source Han Sans J", K, SC, and TC. Then when you use the font, depending on which font name you used the system will change which mapping it applies to the combined set of glyphs. (Hence the choice of font name is the selection mechanism you described.) The second way is that TrueType fonts have a way to build locale settings into the font. I'm less clear on the details here but apparently it's similar to TTC behind the scenes, except that the mappings are associated with locales - so in an app that supports TT locales, even if you select "Foo J" as your font, when the locale was simplified Chinese you'd get the SC glyph. Of course now the selection mechanism is whether the application knows what locale the content is. (And also whether it supports the mechanism - I don't know how widespread this is.) Either way though, in principle you get different glyphs for the same code point, depending on context. Or anyway that's the understanding I took away as a font layperson - happy to be corrected. [0] http://sourceforge.net/projects/source-han-sans.adobe/files/ http://sourceforge.net/projects/source-han-sans.adobe/files/
- yuubi 12y agoThe modest proposal was to extend this approach with all its complications to western alphabetic scripts where possible. Of course nobody wants to do that for obvious reasons, which also apply to the various languages that use Han-derived characters. The extra mechanisms required to work around unified Han remind me of the pre-unicode days when you needed to know what language a text was in to render it.
- Osmium 12y agoSurely there are enough unicode code points for this not be a problem? Can you use the historical character + combining mark (which shows which 'newer' version of the character to use), where the combining mark is ignored if the computer doesn't understand it, and only then it falls back onto guesswork/lang attributes. I don't know if that's a decent solution, but just guesswork doesn't sound like a good idea, because there are bound to be edge cases where it wouldn't work, and then we're back where we started...
- anon4 12y agoThen you break copy-paste. Or maybe we could add locale markers in unicode, then encode the different symbols as <locale><codepoint>. It will only take something like up to 12 bytes per character in UTF-8, no big deal, right?
- jessaustin 12y agoWell done. I admit, this passed over my head on first reading.