5 ms·
Re leveraging UTF-8 for encoding text to be visually small: if you have your conflict-free hash as a binary number, you could do this fairly easily by making an
by infinityio 4y ago
Re leveraging UTF-8 for encoding text to be visually small: if you have your conflict-free hash as a binary number, you could do this fairly easily by making an 'empty' UTF8 char of the right length and replacing all the data bits with your hash, and it would be converted to a potentially invalid, but still transmissible, character for you (if you have to make the characters valid, you may need a lookup table to jump over these)
> My proposed scheme enable to have N * 255 * 255 1 byte characters. With a constant cost per string of 1 byte. This seems revolutionary.
The main issue in this context would be limitations in encoding enough languages - there are too many characters (144k in unicode) to encode in 256 code pages (16k max). Additionally, frequently switching code pages (for example swapping back to ascii to use Latin chars / numbers) would incur a large cost on the size of the string
Also note that some of the 'earlier' unicode symbols that don't fit into ascii are in 2 bytes as well, it's not just 1/3/4!