4 ms·
The beauty of UNIX to me is the interoperability of the text stream What interoperability? Look at the man page for any simple Unix utility (such as `ls`), and
by quanticle 4y ago
The beauty of UNIX to me is the interoperability of the text stream
What interoperability? Look at the man page for any simple Unix utility (such as `ls`), and count up how many of the listed command line flags are there only to structure the text stream for some other program. "Plain text" is just as interoperable as plain binary. "Just use plain text" is the Original Sin of Unix.
The Unix Philosophy, as stated by Peter Salus [1] is
1. Write programs that do one thing and do it well
2. Write programs that work together
3. Write programs to handle text streams because text is a universal interface.
The problem is that, in practice, you can only pick two of those. If you want to write programs that work together, and do so using plain text, then, in addition to doing its ostensible task, each program is going to have to provide a facility to format its text for other programs, and have a parser to read the input that other programs provide, contradicting the dictum to "do one thing and do it well".
If you want programs that do one thing and do it well, and programs that work together, then you have to abandon "plain text", and enforce some kind of common data format that programs are required to read and output. It might be JSON. Or it might be some kind of binary format (like what PowerShell uses). But there has to be some kind of structure that allows programs to interchange data without each program having to deal with the M x N problem of having to deal with every other program's idiosyncratic "plain text" output format.
[1]: https://en.wikipedia.org/wiki/Unix_philosophy https://en.wikipedia.org/wiki/Unix_philosophy
- jimbo9991 4y agoThis was a really insightful take on the Unix philosophy that I hadn't heard before but I intuitively agree with because of all the parsing code I've had to write.
- infogulch 4y agoProgrammers tend to have a disproportionate affinity towards plain text, me included. But this is an intriguing argument so now I'm reconsidering. Maybe plain text is just someone else's unparsed junk.
- tm-guimaraes 4y agoBecause plain text is narrow waist* that is easily debuggable. * https://www.oilshell.org/blog/2022/02/diagrams.html https://www.oilshell.org/blog/2022/02/diagrams.html
- deafpolygon 4y agoThe only thing I really take away from the UNIX philosophy nowadays (I used to be a dyed in the wool fan of UNIX/Linux) is #1) do one thing and do it well. I see #2 as an ideal goal to reach but not always required. And #3 is nowadays untenable for me. If we can agree on an object exchange format (something PowerShell seem to have solved in part), then we can do much much more than relying on text streams.
- quanticle 4y agoThe only thing I really take away from the UNIX philosophy nowadays (I used to be a dyed in the wool fan of UNIX/Linux) is #1) do one thing and do it well. I see #2 as an ideal goal to reach but not always required. If you have a number of programs, each of which does one thing and does it well, those programs will need to exchange data between themselves in order for the overall system to be useful to the user. To go back to the `ls` example I used in my post above: `ls` should just list files. Why should `ls` have anything to do with sorting, when `sort` exists? The reason, as it stands right now, is that `ls`'s plain text output is too much of a pain to parse, and so it's more convenient to build sorting into `ls` itself. If you start with "do one thing and do it well", but ignore interoperability, then your program will inevitably grow additional options and subcommands until it does one thing well, and quite a lot of things mediocrely. Instead of a collection of small sharp specialized tools, you'll end up with, e.g. `find`.
- deafpolygon 4y agoOne thing well, can also be seen as listing directories in many different ways.
- quanticle 4y agoBy that logic `perl` is a tool that does "one thing", where that "one thing" just happens to be "everything".
- chubot 4y agoThe M x N interoperability problem is SOLVED BY building on top of bytes / plain text, not solved by moving AWAY from it! See this section of my follow-up post (to the link below, which is how I found this post): https://www.oilshell.org/blog/2022/03/backlog-arch.html#slogan-text-is-the-only-thing-you-can-agree-on https://www.oilshell.org/blog/2022/03/backlog-arch.html#slog... There have been numerous projects which invent bespoke protocols for interoperability -- I give the examples of PowerShell, Elvish, and nushell. (PowerShell doesn't use any kind of binary format AFAIK. I believe you are literally moving around .NET objects inside the CLR VM, and that representation is meaningless outside the CLR VM. This is crucial because it means that PowerShell must serialize its data structures for interoperability.) As well as various Lisps. The argument is how you interoperate between them. (Honest question -- please let me know.) So ironically, trying to solve the interoperability problem in a smaller context CREATES it again in a bigger context (e.g. between different machines). Bytes and text are fundamental because they reflect how disks and networks fundamentally work, in addition to operating system. That does not mean we shouldn't have higher level layers on top of bytes and text, like JSON, HTML/XML, and TSV/CSV. Those are structured data formats. You generally use parsing libraries for them, instead of writing the parser yourself. Again, all of those formats ARE text, and that's a feature, not a bug!
- quanticle 4y agoThe M x N interoperability problem is SOLVED BY building on top of bytes / plain text, not solved by moving AWAY from it! I never argued that we should move away from plain text. Indeed, one of the examples I cited of an interoperability format, JSON, does exactly that: it takes plain text and adds structure to it to make it easily parseable by machines. Overall, I'm not sure what part of my post you're arguing against. I'm suggesting that instead of having a many different ad-hoc formats, Unix utilities should agree on a few accepted serialization formats and pass information around using those. This would make our shell pipelines less complex and more robust because we wouldn't have to worry about e.g. random spaces or newlines breaking the ad-hoc parsers we write with `grep` and `cut`. It seems like you agree with that, with the caveat that that serialization format should be built on top of plain text. That's fine. We can agree on JSON as a serialization format. It's not ideal, but it's better than the myriad of ad-hoc formats that we have now.