3 ms·
In your quote encoding refers to assigning numbers (code points in Unicode parlance) to characters (I am simplifying here, I know the definition of character in
by brewmarche 4y ago
In your quote encoding refers to assigning numbers (code points in Unicode parlance) to characters (I am simplifying here, I know the definition of character in Unicode is not that easy).
It’s like a catalogue of scripts. We have to extend it when we encounter new scripts that are not catalogued yet (or when we create new emojis)
Converting a byte sequence to a Unicode code point sequence and vice-versa is called transformation format (or more generally an encoding form, but then might not be deterministic) by Unicode (see <https://www.unicode.org/faq/utf_bom.html#gen2 https://www.unicode.org/faq/utf_bom.html#gen2>). Unicode specifies UTF-8, -16 and -32. We do not have to change these formats unless the catalogue hit the limits of 32 bits (not a big problem for UTF-8 but for the other two formats). These formats are already able to encode code points that are not assigned yet.
And the confusion now is that a lot of people call what Unicode calls transformation format (i.e. the byte to code point mapping) encoding as well. The term charset is also used sometimes.
PS:
Note that a goal of Unicode is to be able to accommodate legacy encoding/charsets by having a broad enough catalogue. This is so that these legacy encoding which may come with their own catalogue can be mapped to the Unicode catalogue. So we have control codes (even though not part of any “proper” human script), precomposed letters (there is a code point for à although it could be represented by a + combining `), things like the Greek terminal form of sigma separately encoded, although that could be done in font-rendering (like generally done for Arabic), and a lot more to aid with mapping and roundtrips.
- tzot 4y agoJust a note about the Greek terminal form of sigma: when dealing with Greek numerals, ςʹ (the final form of sigma) is 6 while σʹ (the non-final form of sigma) is 200; they need to be differentiated whatever the font engine decides to render.
- brewmarche 4y agoThanks, good to know. I’ve also realised that determining which form to use in mathematical formulas is maybe not that straightforward. Edit: by which I mean that I’ve only seen the non terminal form so far in maths but it’s hard to write an algorithm that distinguishes between a word ending in sigma and some juxtaposition of Greek variables.