15 ms·
Hex has lower information density, however it is more compressible, particularly if your data is byte aligned. For example the string "aaaaaa" is b64 encoded as
by hackcasual 9y ago
Hex has lower information density, however it is more compressible, particularly if your data is byte aligned. For example the string "aaaaaa" is b64 encoded as "YWFhYWFh" but in hex is "616161616161". On a project I'm on, we switched from base64 to hex for some binary data embedded in JSON, and saw ~20% size reduction, since its always compressed by the web server.
- _wmd 9y agoToday I learned! >>> import random,zlib,codecs >>> sample = random.getrandbits(1024).to_bytes(1024, 'big') * 1024 >>> len(zlib.compress(codecs.encode(sample, 'base64'))) 22083 >>> len(zlib.compress(codecs.encode(sample, 'hex'))) 4326 >>> edit: (25 minutes of digging around the Internet, and I find this): >>> sample = os.urandom(1024) * 1024 >>> len(sample.encode('base64')) 1416501 >>> len(sample.encode('hex')) 2097152 >>> len(sample.encode('ascii85')) 1310720 >>> len(zlib.compress(sample.encode('base64'))) 82169 >>> len(zlib.compress(sample.encode('hex'))) 12537 >>> len(zlib.compress(sample.encode('ascii85'))) 8212 Had to use Python 2.x to make use of the 'hackercodecs' package that implements ascii85, but looks like it's the best of both worlds, assuming its character set suits whatever medium you are transferring over, and decoding it on the other end doesn't require some slow code. final edit: I'm guessing it was an accidental trick of the repetitive input data lining up well. On real data hex still wins out (which should have been obvious in hindsight): >>> sample = ''.join(sorted(open('/usr/share/dict/words').readlines())) >>> len(zlib.compress(sample.encode('ascii85'))) 1320809 >>> len(zlib.compress(sample.encode('base64'))) 1116678 >>> len(zlib.compress(sample.encode('hex'))) 880651
- andreareina 9y agoHuh, my results are drastically different: $ dd if=/dev/urandom bs=512 count=2048 | base64 | gzip | wc ... 4088 31928 1059085 $ dd if=/dev/urandom bs=512 count=2048 | xxd -p | gzip | wc ... 5019 33798 1231268 1 megabyte of random data consistently results in ~1 megabyte of compressed base64 text, ~1.2 megabytes of compressed hex.
- _wmd 9y agoMy data wasn't quite random! It repeats every 1kb, much smaller than zlib's window size (which is I think 16kb)
- nkurz 9y agoA good rule-of-thumb might be that if your results show that you able to consistently compress supposedly random data to less than the size required for just the random binary bits, you should either recheck your numbers, verify your random number generator, or quickly file for a patent!
- jmiserez 9y agoYou should add googling “Shannon” “entropy” and “information theory” to that list ;-)
- Twirrim 9y agoRepeating the exercise using a photograph (https://imgur.com/B4tqkrZ https://imgur.com/B4tqkrZ): $ pv lock-your-screen.png | base64 | gzip | wc 2.1MiB 0:00:00 [ 13MiB/s] [================================>] 100% 8641 48904 2264151 $ pv lock-your-screen.png | xxd -p | gzip | wc 2.1MiB 0:00:00 [4.41MiB/s] [================================>] 100% 10109 49956 2573293 See the same against the jpg version (I've no idea why I have kept both a jpg and png of the same image around, especially given jpg is much better suited format): $ pv lock-your-screen.jpg | base64 | gzip | wc 377KiB 0:00:00 [19.2MiB/s] [================================>] 100% 1420 8373 392796 $ pv lock-your-screen.jpg | xxd -p | gzip | wc 377KiB 0:00:00 [9.36MiB/s] [================================>] 100% 1487 8935 441077 In both cases base64 is both faster and compresses smaller.
- BeeOnRope 9y agoThis test isn't very informative because both .png and .jpg are already compressed formats, with "better than gzip" strength so gzip/deflate isn't going to be able to compress the underlying data. You only see some compression because gzip is just backing out some of the redundancy added by the hex or base64 encoding, and the way the huffman coding works favors base64 slightly. Try with uncompressed data and you'll get a different result. Your speed comparison seems disingenuous: you are benchmarking "xxd", a generalized hex dump tool, against base64, a dedicated base-64 library. I wouldn't expect their speeds to have any interesting relationship with best possible speed of a tuned algorithm. There is little doubt that base-16 encoding is going to be very fast, and trivially vectorizable (in a much simpler way than base-64).
- jmiserez 9y agoTry with (much) more data, theoretically they should even out at some point.
- jwilk 9y agoFWIW, Python (≥ 3.4) has an ascii85 implementation in the standard library: >>> import base64 >>> base64.a85encode(b'spam') b'F)YQ)' https://docs.python.org/3/library/base64.html#base64.a85encode https://docs.python.org/3/library/base64.html#base64.a85enco...
- hackcasual 9y agoBoth hex and ascii85 will align at 1024 byte boundaries, base64 will align at 1024*3 bytes, so you'll end up with a symbol stream something like this: symbol_1, symbol_2, symbol_3, ptr_to_1, ptr_to_2,... ascii85 performs better than hex simply because it's dictionary is shorter. If you compress something with a bit more structure, like "Hamlet", you'll get this: >>> import urllib.request,zlib,codecs,base64 >>> hamlet = urllib.request.urlopen("http://www.gutenberg.org/cache/epub/2265/pg2265.txt").read() >>> len(codecs.encode(hamlet, 'base64')) 248577 >>> len(codecs.encode(hamlet, 'hex')) 368018 >>> len(base64.a85encode(hamlet)) 230012 >>> len(zlib.compress(codecs.encode(hamlet, 'base64'))) 102788 >>> len(zlib.compress(codecs.encode(hamlet, 'hex'))) 88827 >>> len(zlib.compress(base64.a85encode(hamlet))) 121364
- gumby 9y agoIf you’re compressing it anyway why turn it into base64 or ascii hex? The compressed data will be binary anyway, so just compress the input data directly.
- dnet 9y agoOne example are JSON/XML/HTML responses over HTTP: you need the 8-bit ("binary") data in an ASCII format to fit into JSON/XML/HTML, while HTTP provides gzip (or DEFLATE, Brotli, etc.) compression over the whole response if both the client (by including the "Accept-Encoding" header in the request) and the server implementations support it.
- hackcasual 9y agoCorrect, we're sending data points as typed arrays inside a much larger JSON payload.