6 ms·
Why in the world does Unix allow newlines in a filename in the first place? That's just such an obviously brain-damaged idea. There's not a single rational use
by amptorn 5y ago
Why in the world does Unix allow newlines in a filename in the first place? That's just such an obviously brain-damaged idea. There's not a single rational use case for it, yet it breaks nearly every text-based tool you could possibly imagine...
- mistrial9 5y agomy imagined reason is -- because when that terrible day happens, and an important file with some new name, does in fact get a newline in it, the rest of the system now has predictable code paths. Q. Is this related to perl, who knows
- Latty 5y agoUnix filenames are just sequences of bytes, not defined as strings. Most programs parse them as utf-8, but there is nothing mandating that. Obviously that leads to problems.
- ninkendo 5y agoOne pedantic qualification: any byte except 0x2f (`/`) or 0x00. This actually rules out nearly any non-UTF8 character set (besides ASCII.) Quote from Linus, which reminds me of Henry Ford’s “you can have any color you want, so long as it’s black”: > And that one true format is UTF-8. End of story. If you try to talk to the kernel in UCS-2 or anything else, you _will_ fail. https://lore.kernel.org/all/Pine.LNX.4.58.0402141827200.14025@home.osdl.org/ https://lore.kernel.org/all/Pine.LNX.4.58.0402141827200.1402...
- dylan604 5y agoI see your pedantic and raise you: UTF-8 isn't a font though. It's a text encoding.
- marklgr 5y agoString bets not allowed, whatever their encoding ;)
- jcranmer 5y ago> This actually rules out nearly any non-UTF8 character set (besides ASCII.) It doesn't--pretty much any character set that has seen widespread use in the past few decades would be compatible. Any single-byte charsets that are ASCII compatible (such as most Windows CP* sets or the entire ISO-8859-* suite) would work. Most Asiatic charsets (e.g., EUC-JP, Shift-JIS, Big5, GBK) that use variable-width encodings follow the rule that characters in the 0x00-0x7f range are ASCII and subsequent characters in the 0x40-0xff range, and so are themselves compatible as well. So actually the list of notable incompatible charsets is easier to write out: UTF-16, UTF-32, EBCDIC, and ISO-2022-* charsets (which are mode-switching).
- ninkendo 5y agoEh, fair enough. While you’re correct, character sets that are “ascii, but something custom when the high bit is 1” are all just “ascii” to me, in that they are all mutually incompatible for anything other than the first 127 characters, and 8-bit encoding in general has been ubiquitous for nearly as long as ascii has been defined. (Meaning that when most people say “ascii”, they’re actually referring to one of those encodings in practice.) Asiatic character sets are an interesting point though. I wonder how common they were at the time of what Linus wrote…
- jcranmer 5y ago> While you’re correct, character sets that are “ascii, but something custom when the high bit is 1” are all just “ascii” to me Don't call them just "ASCII"--that only serves to confuse people. Call them 8-bit ASCII-compatible charsets if you need a collective noun, but note that they are very different. > (Meaning that when most people say “ascii”, they’re actually referring to one of those encodings in practice.) Having actually worked on charset handling, when most people say "ASCII", they mean "ASCII" and not anything else. If a document is labeled as ASCII, then generally it should be handled as Windows-1252. If a conversion function claims to convert ASCII to something else, and doesn't provide any error mechanism (which it really should), then it usually means ISO-8859-1 aka Latin-1 aka map each byte to the first 256 Unicode characters. But I'd never see, e.g., a KOI8-R document referred to as ASCII, nor anything that claimed to be ASCII assumed to be a KOI8-R document. > Asiatic character sets are an interesting point though. I wonder how common they were at the time of what Linus wrote… https://4.bp.blogspot.com/-O4jXmTm7WWI/Tyw1As8jt7I/AAAAAAAAI9E/nxxi1T21IH4/s1600/unicode.png https://4.bp.blogspot.com/-O4jXmTm7WWI/Tyw1As8jt7I/AAAAAAAAI... At the time he wrote that, the main Asiatic charsets for Chinese and Japanese would have been more common than UTF-8. Maybe Korean as well, although Linus's message is around the time that UTF-8 overtook EUC-KR. In any case, anyone who knew anything about character sets at the time would have been well aware of Asiatic variable-width character sets.
- amptorn 5y ago> Unix filenames are just sequences of bytes, not defined as strings "Write programs to handle text streams, because that is a universal interface except for filenames which are opaque binary"
- jl6 5y agoI can’t think of why you’d ever want a newline in a filename, but it does make for easier reasoning about what characters (or perhaps I should say bytes) could be found in filenames, as opposed to having to remember a long list of exceptions.
- jagrsw 5y ago> yet it breaks nearly every text-based tool you could possibly imagine It breaks badly designed text protocols - some can argue that it's a good idea - "crash early, crash loud" etc. Also if your protocol breaks with newlines, it probably breaks with other non-literals - brackets, quotes, NUL-bytes, control characters, carriage return char, multibyte chars etc etc.
- wutbrodo 5y ago> It breaks badly designed text protocols - some can argue that it's a good idea - "crash early, crash loud" etc This is decisively not a case of "fail loudly", which I agree is generally a good idea. The very first example in the article is one of silent incorrect/ambiguous output, not loud failure.
- tyingq 5y agoIt is odd. Though tools like find have "-print0" for this purpose. And corresponding input flags for xargs, perl, sort, uniq, cut, head, etc, that accept NUL terminated vs newline terminated lists.
- bayindirh 5y agoI'm against limiting the character set allowed for file names. macOS is also in the same boat with Linux, going one step forward and allowing \null terminator even in the filenames. If we're going to limit filenames' character sets, I can offer a simpler solution: Why allow file names? OS should provide a UUID for all files. No names, nothing. We can just write which file is what to another file, noting its UUIDs to sticky notes.
- feldrim 5y agoThis is an old solution to a problem that does not exist. Yes, in that case the file system can be a key-value store. It would eliminate the need for a tree structure. But the tree structure has a meaning: it adds context. The directories are containers of files that adds a semantic abstraction to the files within. https://devblogs.microsoft.com/oldnewthing/20110228-00/?p=11363 https://devblogs.microsoft.com/oldnewthing/20110228-00/?p=11...
- wlib 5y agoWhy do we impose hierarchy so much in file systems? We already allow hard and soft links, so it’s not even a tree anyways. Why not just allow any reference types you want; no name with extensions, but a set of tags. Why not identify files the same way a graph database query identifies nodes?
- feldrim 5y agoSo you propose a graph database for data structures, without the persistence layer provided by the file system, right?
- sitharus 5y agoBecause hierarchical structures and names are easy to explain to most people. macOS has supported tagging for ages, but I’ve never seen it used extensively or as a complete alternative to tree structure.
- dahfizz 5y ago
- marcosdumay 5y agoWhy would Unix go and add random restrictions to filenames? And what text protocol requires you to just insert user data without escaping or re-encoding? That looks badly broken. The kind of broken that will give your entire system to a hacker for encrypting and demanding ransom.
- jlarocco 5y ago> That's just such an obviously brain-damaged idea. Is it, though? "Every character except '/' because it's the directory delimiter" seems pretty straight forward to me... > There's not a single rational use case for it, yet it breaks nearly every text-based tool you could possibly imagine... You don't have a use case, but that doesn't mean nobody else has one. And as far as "text-based tools" go, their developers should RTFM. I'm fairly sure UNIX existed before almost all of them, and it's accepted new lines all along.
- dzaima 5y agoWhy not also, while at it, disallow spaces too? They can very easily cause problems too, if you split by spaces instead of newlines. Quotes and backslashes obviously are also bad. How about all of non-ASCII unicode? That'd break all code assuming character count equals byte count, and can probably cause buffer overflows when people count correctly. Any characters you disallow still allows people to fail on some other character. Sure, it'd decrease the likelihood of messing things up by some amount, but that's a half-assed solution at best, and would make people check for mistakes less at worst. Imagine if intel fixed the pentium FDIV bug by only fixing 30% of the wrong results.
- kroltan 5y agoNo, write your software properly. Assuming anything at all about file names is how we get to silly things like Windows' "CON" or whatever restrictions.