4 ms·
I compressed "I am going to work outside today," then put the compressed output in Google Translate. Google translated the Chinese characters back to English as
by starpilot 6y ago
I compressed "I am going to work outside today," then put the compressed output in Google Translate. Google translated the Chinese characters back to English as "raccoon."
- dhosek 6y agoI think the Chinese text that comes out confuses Google translate. I took the whole first sentence of Hamlet's soliloquy which compressed to 䮛趁䌆뺜㞵蹧泔됛姞音逎贊 and plugged that into Google Translate. It came back with "Commendation." The reverse translation is 表彰
- dhosek 6y agoI hit the reverse button again, and got "recognition" which translated back to 承认 which finally got into a closed loop to recognition and back to the same Chinese text.
- james412 6y agoIt's not Chinese text, it's an arithmetic-coded stream of bits mapped so the bits fall within the range of some codepoints. It's basically a variant of base64 except for Unicode. (Side note: aren't these codepoints very expensive to encode in UTF-8? It seems there must be a lower-valued range more suited to it)
- willcipriano 6y agoIt's probably similar to this: https://pieroxy.net/blog/pages/lz-string/index.html https://pieroxy.net/blog/pages/lz-string/index.html Check the 'How does it look?' section.
- greenshackle2 6y agoYeah I don't understand why it's using CJK, the page claims: > each compressed character holds 15 data bits by using the CJK and the Hangul Syllables unicode ranges. In UTF-8 these characters take 3 bytes each. Which makes it less space efficient than base64 (60% overhead vs 33% overhead). The CJK/Hangul scheme has more information per character but I'm not sure where that matters.
- jkhdigital 6y agoIt's so that the output uses printable characters, that's all. The raw output would actually just be random bits, or at least something approximating random bits if GPT-2 is as good as we hope it is.
- greenshackle2 6y agoYes, I understand. But base64 is the bog standard solution for encoding arbitrary binary data as printable characters. So I'm just wondering why you would use something more obscure, less space efficient, and not ascii compatible.
- im3w1l 6y agoBecause it will impress uneducated people with how much smaller (in terms of screen real estate) the resulting message is. EDIT: Ohhh, I know, and because twitter cares about characters. So you can use this to put essays into tweets.
- toast0 6y agoThe page for base32768 has some efficiency charts for different binary to text encodings on top of different UTF encodings, as well as how many bytes you can use them to stuff in a tweet. Depends on where you're going to house the data, I guess. https://github.com/qntm/base32768 https://github.com/qntm/base32768
- infogulch 6y agoIn addition to being 94% efficient in UTF-16 (!), this reveals some additional reasons why one might want to optimize for number of characters: fitting as many bytes as possible into a tweet which is bounded in the number of characters not bytes.
- dheera 6y agoThis is Chinese characters mixed with Korean characters and this is pretty much never done by humans. It is analogous to mixing English and Heiroglyphics and typing out some gibberish with both. The author might as well have included the rest of the Unicode range including Arabic, Emoji, and math symbols.
- gHosts 6y agoFor more fun enter "I am going to work outside today" compress, delete the second character and decompress, the result is... "I know what you're thinking –"