39 ms·
Actually, my proposed encoding only needs to add 64 to the code point (or subtract 64 when encoding a code point), not 128. That in mind, ♥ is encoded as follo
by strenholme 14d ago
Actually, my proposed encoding only needs to add 64 to the code point (or subtract 64 when encoding a code point), not 128. That in mind, ♥ is encoded as follows:
♥ → 0b0010_0110_0110_0101 → 0b0010_0110_0010_0101 (subtract 64) → 0b10_00_0010 0b10_0110_00 0b11_10_0101
Code point 128 is encoded as follows:
128 → 0b1000_0000 → 0b0100_0000 (subtract 64) → 0b10_0000_01 0b11_00_0000
C99 can do up to over 10 bytes long (uint64_t) without a bignum library since 10 bytes gives us 60 bits. C23 gives us a bignum library so the sky’s the limit.
Again, I only need 7 bits to represent every character the font for my blog has.
- strenholme 13d agoAnother extension to this encoding: If we have an 0b11xx_xxxx byte which isn’t proceeded by a 0b10xx_xxxx byte, the numeric value of the byte is its corresponding codepoint. This gives us all of the accented letters western European languages use, allowing us to represent all of ASCII and all western European letters with only one byte. This also means that the top half of ISO 8859-1 will have two representations with this encoding, but since they are letters, with the only symbols being × (multiplication) and ÷ (division), this should not be a security risk, unless one programs in Raku (which, ugh, gives meta significance to non-ASCII Unicode symbols).