4 ms·
In b64, you're going to want to pick a short set of characters that can be encoded clearly and simply, i.e. ascii chars. And the shorter the better, i.e. even i
by edblarney 10y ago
In b64, you're going to want to pick a short set of characters that can be encoded clearly and simply, i.e. ascii chars. And the shorter the better, i.e. even in 6 bits.
You take the first six bits of your binary data and convert to some ascii char mapping. And then the the next 6 bits and so on.
You can't do that with unicode, and it wouldn't make sense for any other encoding standard.
They are completely different things.
- Moru 10y agoMabe srean wants a more efficient way of base64 a binary file into some text-only media like email. This depends totally on where you want to put it. If it's an email you are pretty stuck with base64, if it's a string that nothing else touches, you can use the binary data directly (eg: iso-8859) :)
- srean 10y agoThat's indeed right. B85 is already a little more efficient than B64, but was wondering if one could abuse Unicode for this. Its a really silly, stupid situation I need this for, exchanging data with Python and I can only use Unicode strings.
- masklinn 10y ago> That's indeed right. B85 is already a little more efficient than B64, but was wondering if one could abuse Unicode for this. Sure, kinda: https://github.com/pfrazee/base-emoji https://github.com/pfrazee/base-emoji (it's a base256 using emoji), but then you still need to encode that text, which is going to require 4 bytes per symbol, so I'm not sure you're going to get any actual gain over B64/B85 in the end. There's also the option of using a subset of the U+0100~U+07FF range (though it contains a diacticial block which may not be ideal) as it encodes to 2 bytes in both UTF-8 and UTF-16 (though there are diacritics in these blocks, and some of the codepoints are reserved but not allocated so…).