12 ms·
Why did base64 win against uuencode?
- t-3 3y agoBase64 is very bizarre in general. Why did they use such a weird pattern of symbols instead of a contiguous section, or at least segments ordered from low->high (on that note, ASCII is also quite strange, I'm guessing due to some backwards compatibility idiocy that seemed like it made sense at some point (or maybe changing case was super important to a lot of workloads or something, making a compelling reason to fuck over the future in favor of optimisation now))?
- eesmith 3y agoThe comments point out conversion issues with EBCDIC. You can't use ASCII characters like @ which are not in EBCDIC. https://datatracker.ietf.org/doc/html/rfc2045#section-6.8 https://datatracker.ietf.org/doc/html/rfc2045#section-6.8 says: This subset has the important property that it is represented identically in all versions of ISO 646, including US-ASCII, and all characters in the subset are also represented identically in all versions of EBCDIC. Other popular encodings, such as the encoding used by the uuencode utility, Macintosh binhex 4.0 [RFC-1741], and the base85 encoding specified as part of Level 2 PostScript, do not share these properties, and thus do not fulfill the portability requirements a binary transport encoding for mail must meet. If you want to learn why ASCII is the way it is, try "The Evolution of Character Codes, 1874-1968" at https://archive.org/details/enf-ascii/mode/2up https://archive.org/details/enf-ascii/mode/2up by Eric Fischer (an HN'er). My reading is contiguous A-Z was meant for better compatibility with 6-bit use.
- mikecoles 3y agoI thought the ASCII upper-case <-> lower-case being a bit operation as being clever.
- duskwuff 3y agoIn the context of a terminal, the Control key is also a bitwise operation. Shifted numerals were nearly a bitwise operation as well, but we didn't end up using that keyboard layout.
- throw0101a 3y ago> I thought the ASCII upper-case <-> lower-case being a bit operation as being clever. From "Things Every Hacker Once Knew" (2017), has an entire section on ASCII and the clever bit-fiddling that occurs: * http://www.catb.org/~esr/faqs/things-every-hacker-once-knew/#_ascii http://www.catb.org/~esr/faqs/things-every-hacker-once-knew/... * Discussion from ~2 months ago: https://news.ycombinator.com/item?id=37701117 https://news.ycombinator.com/item?id=37701117
- eesmith 3y agoYes, though in principle you could interleave AaBbCc and so on, which would also be a single bit difference, and the naive collation would be more like that people expect. The design considerations at https://ia800606.us.archive.org/17/items/enf-ascii-1972-1975/Image070917152640_text.pdf https://ia800606.us.archive.org/17/items/enf-ascii-1972-1975... show that 6-bit support was more important than naive collation support: > A6.4 It is expected that devices having the capability of printing only 64 graphic symbols will continue to be important. It may be desirable to arrange these devices to print one symbol for the bit pattern of both upper and lower case of a given alphabetic letter. To facilitate this, there should be a single-bit difference between the upper and lower case representations of any given letter. Combined with the requirement that a given case of the alphabet be contiguous, this dictated the assignment of the alphabet, as shown in columns 4 through 7. I just found and skimmed Bob Bemer's "A Story of ASCII", which includes personal recollections of the history. It seems that the 6-bit subset was firmed up first. From https://archive.org/details/ascii-bemer/page/n17/mode/2up?q=lower https://archive.org/details/ascii-bemer/page/n17/mode/2up?q=... : > This is reflected in the set I proposed to X3 on 1961 September 18 (Table 3, column 3), and these three characters remained in the set from that time on. The lower case alphabet was also shown, but for some time this was resisted, lest the communications people find a need for more than the two columns then allocated for control functions. but serious discussion of lower case wasn't taken up until later. From https://archive.org/details/ascii-bemer/page/n25/mode/2up?q=lower https://archive.org/details/ascii-bemer/page/n25/mode/2up?q=... : > ISO/TC97/SC2 held its next meeting in 1963 October, at which time it was decided to add the lower case alphabet. and at https://archive.org/details/ascii-bemer/page/n27/mode/2up?q=lower https://archive.org/details/ascii-bemer/page/n27/mode/2up?q=... : > At the 1963 May meeting in Geneva, CCITT endorsed the principle of the 7-bit code for any new telegraph alphabet, and expressed general but preliminary agreement with the ISO work. It further requested the placement of the lower case alphabet in the unassigned area. Bemer did not like interleaving lower- and upper-case. From https://archive.org/details/ascii-bemer/page/n5/mode/2up?q=lower https://archive.org/details/ascii-bemer/page/n5/mode/2up?q=l... : > I had a great opportunity to start on the standards road when invited by Dr. Werner Buchholz to do the main design of the 120-character set [9,24] for the Stretch computer (the IBM 7030). I had help, but the mistakes are all mine (such as the interspersal of the upper and lower case alphabets). ... > he didn't make the same mistake I made for STRETCH by interspersing both cases of the alphabet!
- layer8 3y agoThe original specification is in RFC 989 [0] from 1987, called “Printable Encoding”, where it explains “The bits resulting from the encryption operation are encoded into characters which are universally representable at all sites, though not necessarily with the same bit patterns […] each group of 6 bits is used as an index into an array of 64 printable characters; the character referenced by the index is placed in the output string. These characters, identified in Table 1, are selected so as to be universally representable, and the set excludes characters with particular significance to SMTP (e.g., ".", "<CR>", "<LF>").” Using the array-indexing method, the noncontiguity of the characters doesn’t matter, and the processing is also independent of the character encoding (e.g. works exactly the same way in EBCDIC). [0] https://www.rfc-editor.org/rfc/rfc989.html#page-9 https://www.rfc-editor.org/rfc/rfc989.html#page-9
- gumby 3y agoLook at ASCII mapped out with four bits across and four bits down and the logic may suddenly snap into place. Also remember that it was implemented by mechanical printing terminals.
- jibal 3y agoBase64 and ASCII both made perfect sense in terms of their requirements, and the future, while not fully anticipated at the time, is doing just fine, with ASCII being now incorporated into largely future-proof UTF-8. Considerably stranger in regard to contiguity was EBCDIC, but it too made sense in terms of its technological requirements, which centered around Hollerith punch cards. https://en.wikipedia.org/wiki/EBCDIC https://en.wikipedia.org/wiki/EBCDIC There are numerous other examples where a lack of knowledge of the technological landscape of the past leads some people to project unwarranted assumptions of incompetence onto the engineers who lived under those constraints. (Hmmm ... perhaps I should have read this person's profile before commenting.)
- deleted 3y ago[deleted]
- t-3 3y agoI never questioned the competence of past engineers, I question the use of backwards compatibility. Hardware has advanced, but software depends on standards and conventions formulated for far less capable hardware, and that's a problem. The efficiency of string processing/generation is hugely important in terms of global energy consumption. A simple and extremely common int->hex string conversion takes twice as many instructions as it would if ASCII was optimized for computability. Bounds-checking for the English alphabet requires either an upfront normalization or twice the checking, so 50-100% more instructions for that. There are also inconsistencies like front and back braces/(angle)brackets/parens not being convertible like the alphabet is. [({< <-> >})] would have been just as or more useful than the alphabet being convertible and saved a few instructions in common parsing loops.
- eesmith 3y ago> takes twice as many instructions What is your preferred system? How does it affect other needs, like collation, or testing if something is upper-case vs. lower-case, or ease of supporting case-insensitivity? Have you measured the performance difference? https://johnnylee-sde.github.io/Fast-unsigned-integer-to-hex-string/ https://johnnylee-sde.github.io/Fast-unsigned-integer-to-hex... shows a branchless UlongToHexString which is essentially as fast as a lookup table and faster than the "naive" implementation. > Bounds-checking for the English alphabet In the following it goes from 2 assembly instructions to three: int is_letter(char c) { c |= 0x20; // normalize to lowercase return ('a' <= c) && (c <= 'z'); } Yes, that's 50% more assembly, to add a single bit-wise or, when testing a single character. But, seriously, when is this useful? English words include an apostrophe, names like the English author Brontë use diacritics, and æ is still (rarely) used, like in the "Endowed Chair for Orthopædic Investigation" at https://orthop.washington.edu/research/ourlabs/collagen/people-collagen-biology-and-genetic-disorders-lab.html https://orthop.washington.edu/research/ourlabs/collagen/peop... . And when testing multiple characters at a time, there are clever optimizations like those used in UlongToHexString. SIMD within a register (SWAR) is quite powerful, eg, 8 characters could be or'ed at once in 64 bits, and of course the CPU can do a lot of work to pipeline things, so 50% more single-clock-tick instructions does not mean %50 more work. > like front and back braces/(angle)brackets/parens not being convertible I have never needed that operation. Why do you need it? Usually when I find a "(" I know I need a ")", and if I also allow a "[" then I need an if-statement anyway since A(8) and A[8] are different things, and both paths implicitly know what to expect. > and saved a few instructions in common parsing loops. Parsing needs to know what specific character comes next, and they are very rarely limited to only those characters. The ones I've looked use a DFA, eg, via a switch statement or lookup table. I can't figure out what advantage there is to that ordering, that is, I can't see why there would be any overall savings. Especially in a language like C++ with > and >> and >>= and A<B<int>> and -> where only some of them are balanced.
- aap_ 3y agoWhat do you find strange about ASCII?
- pravus 3y ago> I'm guessing due to some backwards compatibility idiocy that seemed like it made sense at some point ... > ... making a compelling reason to fuck over the future in favor of optimisation now > I never questioned the competence of past engineers False just based on your opening volley of toxic spew. Backwards compatibility is an engineering decision and it was made by very competent people to interoperate with a large number of systems. The future has never been fucked over. You seem to not understand how ASCII is encoded. It is primarily based on bit-groups where the numeric ranges for character groupings can be easily determined using very simple (and fast) bit-wise operations. All of the basic C functions to test single-byte characters such as `isalpha()`, `isdigit()`, `islower()`, `isupper()`, etc. use this fact. You can then optimize these into grouped instructions and pipeline them. Pull up `man ascii` and pay attention to the hex encodings at the start of all the major symbol groups. This is still useful today! No, the biggest fuckage of the internet age has been Unicode which absolutely destroys this mapping. We no longer have any semblance of a 1:1 translation between any set of input bytes and any other set of character attributes. And this is just required to get simple language idioms correct. The best you can do is use bit-groupings to determine encoding errors (ala UTF-8) or stick with a larger translation table that includes surrogates (UTF-16, UTF-32, etc). They will all suffer the same "performance" problem called the "real world".
- smudgy 3y agoYou know, when I first got into binary encoding into text I asked myself this very question but never put any effort into looking it up. Now, 25+ years later, I have some answers - thanks!
- DaiPlusPlus 3y agoOn a related note, I'm getting flashbacks to being on the web in the late-1990s, back when "Downloads!" was a reason to visit a particular website; and noticing that Windows users like myself could just download-and-run an .exe file, while the same downloads for Mactintosh users would be a BinHex file that'd also be much larger than the Windows equivalent - and this wasn't over FTP or Telnet, but an in-browser HTTP download, just like today. Can anyone explain why BinHex remained "popular" in online Mac communities through to the early 2000s? Why couldn't Macs download "real" binary files back then?
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- ksherlock 3y agoClassic Macintosh files were basically 2 separate files with the same name (data fork and resource fork). Additionally, there was important meta data (Finder Info, most importantly the file type, creator type). Since other file systems couldn't handle forks or finder info, it had to be encapsulated in some other format like binhex, macbinary, applesingle, or stuffit. The other 3 were binary so they would have been smaller. Why not them... shrug
- ksherlock 3y agoI wasn't a macintosh user back in the day but for the file archives I frequented (apple II), sometimes files were in BINSCII which was a similar text encoding. The advantage being that they could be emailed inline, posted to usenet, didn't require an 8-bit connection (important back in the 80s), and could be transferred by screen scraping if there wasn't a better alternative.
- devilbunny 3y agoSo this is really random, but Kermit actually could route around your 7-bit issues. Everyone remembers it as a godawful protocol choice, because terminal programs usually implemented the most basic version of the protocol, but if configured correctly it was on par with or better than Zmodem.
- ggm 3y agoHaving lived through the transition, I can say personally it comes down to "packaging" -if MIME had adopted UUENCODE format, I probably would have used it but as materials emerged to me which depended on base64 decode, it became compelling to use it. Once it was ubiquitously available in e.g ssl, it became trivial to decode a base64 encoded thing, no matter what. Not all systems had a functioning uudecode all the time. DOS for instance, you had to find one. If you're given base64 content, you install a base64 encode/decode package and then its what you have. There was also an extended period of time where people did uux much as they did shar: both of which are inviting somebody else's hands into your execution state and filestore. We were also obsessed with efficiency. base64 was "sold" as denser encoding. I can't say if it was true overall, but just as we discussed lempel-zif and gzip tuning on usenet news, we discussed uuencode/base64 and other text wrapping. Ned Freed, Nathaniel Borenstein, Patrik Falstrom and Robert Elz amongst others come to mind as people who worked on the baseXX encoding and discussed this on the lists at the time. Other alphabets were discussed. uu* was the product of Mike Lesk a decade before, who was a lot quieter on the lists: He'd moved into different circles, was doing other things and not really that interested in the chatter around line encoding issues.
- mjevans 3y agoAfter a given point usenet was nearly 8-bit clean, and thus https://en.wikipedia.org/wiki/YEnc https://en.wikipedia.org/wiki/YEnc was also developed to convolve all the octets (I + 42 (decimal)) and escape the results that happened to still match reserved characters (CR, LF, 0x0, = (yEnc escape)) - it seems that if the result character was among that set, then = was output and new output determined by O = (I+64) % 256 instead.
- ekidd 3y agoOne reason that uuencode lost out to Base64 was that uuencode used spaces in its encoding. It was fairly common for Internet protocols in those days to mess with whitespace, so it was often necessary to patch up corrupted uuencode files by hand. Base64, on the other hand, was carefully designed to survive everything from whitespace corruption to being passed through non-ASCII character sets. And then it became widely used as part of MIME.
- jjeaff 3y agoand yet, Internet protocols (http, at least) don't play well with equal signs which are part of base64, sometimes. That little issue has caused lots of intermittent bugs for me over the years, either from forgetting to urlencode it or not urldecoding it at the right time.
- Vt71fcAqt7 3y agoAnd now we can have whitespace in url queries but we are still using %20 everywhere because "that's standard"...
- CydeWeys 3y agoTry copy-pasting a link that has actual whitespace in its URL queries and see if it gets linkified correctly. Just because you can doesn't mean you should! A space is like the one delimiter that is applicable for separating out URLs from the context of a larger blob of text.
- thaumasiotes 3y agoBrowsers will often display %20 as a space, but that's not the same thing as spaces being legal within URLs.
- Vt71fcAqt7 3y agoYou are right. Seems firefox displays %20 as whitespace and converts whitespace to %20 when you use it. Chrome displays it as %20 but still converts whitespace to %20 if you try to use it.
- quickthrower2 3y agoI first met base64 in ASP.NET viewstate.
- Aardwolf 3y agoA thing I wonder: why is using = padding required in the most common base64 variant? It's redundant since this info can be fully inferred from the length of the stream. Even for concatenations it is not necessary to require it, since you must still know the length of each sub stream (and = does not always appear so is not a separator). There's no way that using the = instead of per-byte length-checking gains any speed, since to prevent reading out of bounds you must check the per byte length anyway, you can't trust input to be a multiple of 4 length. It could only make sense if it's somehow required to read 4 bytes at once, and you can't possibly read less, but what platform is such?
- remram 3y ago> Even for concatenations it is not necessary to require it, since you must still know the length of each sub stream I'm not sure I understand this part. You can decode aGVsbG8=IHdvcmxk, what do you need to know?
- Aardwolf 3y agoThe = does not appear if the base64 data is a multiple of 4 length. So you wouldn't know if aGVsbG8I is one or two streams. The = is not a separator, only padding to make the base64 stream a multiple of 4 length for some reason. I only mentioned the concatenation because Wikipedia claims this use case requires padding while in reality it doesn't.
- lifthrasiir 3y agoBase64 doesn't have a concept of "stream". Conceptually base64-encoded string with padding is a concatenation of fragments that are always 4 bytes long but can encode one to three bytes. Concatenating two base64-encoded strings with padding therefore don't destroy fragment structures and can be decoded into a byte sequence that is a concatenation of two original input sequences. Without padding, fragments can be also 2 or 3 bytes and short fragments are not distinguishable from long fragments, so the concatenation will destroy fragment structures.
- 38 3y agohttps://wikipedia.org/wiki/Binary-to-text_encoding https://wikipedia.org/wiki/Binary-to-text_encoding
- whoopdedo 3y agoNot listed was a clever encoding for MS-DOS files, XXBUG[1]. DOS had a rudimentary debugger and memory editor. (It even stuck around all the way to Windows XP but didn't survive the transition to 64-bit.) Because it had the ability to write to disk you could convert any file to hexadecimal bytes and sprinkle some control commands about to create a script for DEBUG.EXE. The text-encoded file could then be sent anywhere without needing to download a decoder program first. [1] http://justsolve.archiveteam.org/wiki/XXBUG http://justsolve.archiveteam.org/wiki/XXBUG
- bnjf 3y agobut uuencode makes for fun: https://github.com/bnjf/compre.sh https://github.com/bnjf/compre.sh
- Dwedit 3y agoThere is still one sort-of efficient way of embedding binary content in an HTML file. You must save the file as UTF-16. A Javscript string from a UTF-16 HTML file can contain anything except these: \0, \r, \n, \\, ", and unmatched surrogate pairs 0xD800-0xDFFF. If you escape any disallowed character in the usual way for a string ("\0", "\r", "\n", "\\", "\"", "\uD800") then there is no decoding process, all the data in the string will be correct. If you throw data that is compressed in there, you're unlikely to get very many zeroes, so you can just hope that there aren't too many unmatched surrogate pairs in your binary data, because those get inflated to 6 times their size. Note that this operates on 16-bit values. In order to see a null, \r, \n, \\ and ", the most significant byte must also be zero, and in order for your data to contain a surrogate pair, you're looking at the two bytes taken together. When the data is compressed, the patterns are less likely.
- nibbula 3y agoAscii85 survived by hiding in popular bloat.