3 ms·
Sorry if I am totally wrong but could you optimise UTF-8 by using some kind of character to say this is a UTF-8 character. For example: AÉBéC would give 0x41F
by geophertz 8y ago
Sorry if I am totally wrong but could you optimise UTF-8 by using some kind of character to say this is a UTF-8 character.
For example:
AÉBéC would give
0x41FF42FF42
where FF would mean, refer to another table with the index of the UTF8 char like the following:
---------------
|------|------|
|0 |1 |
|------|------|
|0xC389|0xC3A9|
|-------------|
This would make random access fast however will also increase overhead at other places.
Also IIRC, UTF8 doesn't use values 0x80 <= unused_utf8_values <= 0xFF. Value between 0x80 and 0xFF could be used to refer to an index of the table. eg 0x80 = index 0, 0x81 = index 1 ... 0xFE = index 126 and 0xFF = refer to other table.
Regardless of values over 0x80. The index in the special table will still have to be kept when iterating over it for find and stuff.
EDIT: table formatting
- Animats 8y agoYou can recognize a "UTF-8 character" in a UTF-8 string easily. You can get the next character, and you can back up to a previous character, all unambiguously. The encoding is kind of neat. See Wikipedia.
- chronial 8y agoRemember that e.g. any chinese text will contain only non-ascii characters.
- deathanatos 8y agoThe "0" "1" on the top of your table make me curious how I translate a 0xff octet into a position in the table. As it's written, it seems like I need to know whether the 0xff I'm looking at is the first, second, etc., which would require scanning the string to figure that out, which would defeat the lookup table entirely. Perhaps we could fix this by storing the index of the 0xff, mapped to codepoint: index => codepoint 1 => 0xc389 3 => 0xc3a9 This would require a binary search; we could change that to a hash table, to make it O(1) w/o too much trouble. However… This degrades pretty badly for strings that don't contain any/much ASCII (Greek, Russian, any Asian language, Indian, …); each of these table indexes is also presumably sizeof(size_t) (for the index) + 4 (for the codepoint), so 12B, probably 16 bytes/slot (after padding). That's 16B per codepoint. Perhaps some more could be squeezed out w/ some sort of variable length deal, but we're already in over our heads, I think. > UTF8 doesn't use values 0x80 <= unused_utf8_values <= 0xFF. UTF-8 uses all octet values up to an including 0xf5, except 0xc0 and 0xc1 (see the table[1]), so there's a few values, but not many. [1]: https://en.wikipedia.org/wiki/UTF-8#Codepage_layout https://en.wikipedia.org/wiki/UTF-8#Codepage_layout