3 ms·
Yes. Unicode is a method of converting between code points and characters. A character is a symbol used in a human language or other written communication. A co
by inklesspen 17y ago
Yes. Unicode is a method of converting between code points and characters. A character is a symbol used in a human language or other written communication. A code point is a number, commonly written as "U+<decimal integer expression of the number>".
There are many different ways to actually store these sequences of code points in a computer. UTF-32 is one of those ways. It takes the decimal integer, coverts it to a 32-bit binary integer, and then splits that number into four 8-bit bytes. As the book says, there are problems with space usage -- in ordinary English text, all the code points will be U+127 or less, which leads to a lot of zero bytes taking up space. In addition to the waste of space, zero bytes can cause problems in C, since they're the symbol for the end of the string. So people invented other 'encodings' to convert Unicode code-points into bytes. UTF-16, UTF-8, etc.