35 ms·
The key flaw of UNIX philosophy is destructuring deserialization and reserialization based on lines, necessitating all manner of argument escaping and field del
by bigtrakzapzap 7y ago
The key flaw of UNIX philosophy is destructuring deserialization and reserialization based on lines, necessitating all manner of argument escaping and field delimiters, when pipelines should be streams of typed messages that encapsulate data in typed fields. Logs especially (logging to files is a terrible idea, because it creates log rotation headaches and each program requires a log parser because of the loss of structured information) also. Line-oriented pipeline processing is fundamentally too simple. Settling on a common data simple/universal format robust enough for all purposes, including very large data sets, to exchange between programs and is just complicated enough to eliminate escapement and delimiter headaches without throwing away flexibility (by a new/refined set of processing tools/commands) is key.
- whatshisface 7y agoIt's not so much a flaw as a tradeoff. Because it's so easy to produce lines of text, virtually every program supports it, even programs written by hackers who aren't interested in following OS communication standards.
- bigtrakzapzap 7y agoNo, that's your opinion, not mine. It's a flaw because it creates escaping magic, delimiters that are in-band (spaces) and every program has to know how to parse the output of every other program, wasting human time and processing time. Structured communication, say like an open standard for a schema-less, self-describing, flexible de/serialization like protobufs/msgpack, would be far superior, more usable, more efficient and simple but no too simple to process streams of data with structure and programmability already there. Being able to dump structured information out of an exception directly into a log, and then from a log into a database, without any loss of information or extraneous log parsing, is a clear win. Or from a command as simple as listing files ("ls") into a database or into any other tool or program. Outputing only line-oriented strings is just throwing away type information, and creates more work for everyone else, even more so than continuing to stay with lines processing tools.
- sk5t 7y agoSo, if you're not a fan of the UNIX philosophy, maybe check out Powershell. Or take a look at WMI and DCOM in Windows. Eschew shell scripts in favor of strongly-typed programs that process strongly-typed binary formats, or XML, or whatever. The alternatives are out there. "Worse is better" isn't for everyone.
- reilly3000 7y agoIt’s worth noting that Powershell is available on Linux. Objects are pretty cool. https://docs.microsoft.com/en-us/powershell/scripting/install/installing-powershell-core-on-linux?view=powershell-6 https://docs.microsoft.com/en-us/powershell/scripting/instal...
- twblalock 7y agoIt works pretty well on the Mac now, too.
- JohnBooty 7y agoNushell does not seem like a violation of the unix philosophy, or at least the version of it that I like best. "Write programs that do one thing and do it well. Write programs to work together. Write programs to handle text streams, because that is a universal interface." Perhaps I'm wrong, isn't nushell simply adding one more dimension? Instead of a one-dimensional array of lines, the standard is now... a two-dimensional array of tabular data. Perhaps it is not strictly a "text stream" but this does not seem to me to be a violation of the spirit of the message. Simple line-based text streams are clearly inadequate IMO. 99.99% of all Unix programs output multiple data fields on one line, and to do anything useful at all with them in a programmatic fashion you wind up needing to hack together some way of parsing out those multiple bits of data on every line. maybe check out Powershell I left Windows right around the time PS became popular, so I never really worked with it. It seems like overkill for most/all things I'd ever want to do. Powershell objects seem very very powerful, but they seemed like too much. Nushell seems like a nice compromise. Avoids the overkill functionality of Powershell.
- kerng 7y agoYes, I agree that it's a tradeoff. Although that tradeoff was made when memory was sparse and expensive, a modern shell should certainly be more structured - e.g. like PowerShell or Nushell here.
- drudru11 7y agoHow should logging be done?
- fouc 7y agoI don't think I'd go as far as typed messages as that doesn't strike me as a simple or light-weight format. Something like JSON is almost good enough, but perhaps something even simpler/lighter would be ideal here.
- mycall 7y agoJSON should have included comments from day 1. Big oops.
- readams 7y agoIt used to. They were removed later
- cben 7y agoComments are great for human-edited config files, but not a clear win for an stdout/in format. The trouble is, comments are normally defined as having no effect, not part of the data model. But iff the comments contain useful information, how do you pick it out using the next tool? Imagine having to extend `jq` with operators to select "comment on line before key Foo" etc... And if it is extractable, are they still comments? How do you even preserve comments from input to output, in tools like grep, sort, etc. XML did make comments part of its "dataset". That is, conforming parsers expose them as part of the data. Similarly for whitespace (though it's only meaningful in rare cases like pre tag but that's up to the tool interpreting data), and some other things that "should not matter" like abbreviated prefixes used for XML namespaces). This does allow round-tripping, but complicates all tools, most importantly by disallowing assumptions that some aspect never matters, e.g. that's it's safe re-indent. I'd argue the depest reason JSON won over XML is that XML's dataset was so damn complicated.
- colordrops 7y agois "destructuring deserialization and reserialization based on lines" really an aspect of the UNIX philosophy or just an implementation detail? I thought it was more about doing one thing and doing it well [1]. It could be argued that nushell follows the UNIX philosophy. [1] https://en.wikipedia.org/wiki/Unix_philosophy https://en.wikipedia.org/wiki/Unix_philosophy
- jen20 7y ago“This is the Unix philosophy: Write programs that do one thing and do it well. Write programs to work together. Write programs to handle text streams, because that is a universal interface.” Doug McIlroy (2003). The Art of Unix Programming: Basics of the Unix Philosophy That said, I’m looking forward to testing our nushell!
- signa11 7y ago> The key flaw of UNIX philosophy is destructuring deserialization and reserialization based on lines not so sure about your assertion there eugene :o) interfacing with funk formatted lines of text is waay more easier because (imho) of the 'wysiayg' principle, and encourages, for lack of a better term, 'tinkering' > Settling on a common data simple/universal format robust enough for all purposes... output from 'ls, gcc, yacc/bison, lex/flex, grep, ...' should all follow the same fmt, sure...the good thing about standards is that we don't have enough of them to choose from, f.e. xml, sgml, s-expressions, json, csv, yaml, ... having said that, when dealing with copious volumes of data, structure doesn't hurt, but in such cases, my guess is, there are only a handful of data-producers, and data-consumers for automatic consumption, and are all tightly controlled.
- barrkel 7y agoIt was a decision that lead to success. Before Unix most filesystems were record oriented. Unix's philosophy of treating everything as a stream of char was different and very successful.
- AnonymousPlanet 7y agoI can understand your frustration. However, the unstructuring of data and keeping a single form of communication is what has kept UNIX around for so long and why Linux hasn't collapsed under its own weight long ago. I was voicing your opinions exactly when I was new to the professional computing world. Over time I saw a lot of structured data schemes come and go and they all fall down with this: inflexibility and improper implementations. Where do you think does the structure inside the data come from? You need to come up with a standard. Now you need everyone to stick to that standard. You need to parse the entire data or at least most of it to get to the information you need, finding new security footguns on the way. Soon you will realise you need to extend the standard. Now you have not only conflicting implementations (because noone ever gets them completely right) but conflicting versions. And this needs to be done for every single fucking little filter and every single language they are written in. Take a look at the various implementations of DER and ASN.1. The standard seems simple at first glance, but I haven't seen a single implementation that wasn't wrong, incomplete or buggy. Most of them are all of that. And DER is a very old standard that people should have understood in the meantime. In order to get at least a bit of sanity in all of this, you need a central source of The Truth wrt. your standard. Who is that going to be for Linux and the BSDs and for macOS? Linus? Microsoft? Apple? The ISO standards board? You? And all of this is equally true for log data. I'm okay with tossing backwards compatibility over board. But not in favour of the horrible nightmare some people here seem to propose.
- yourapostasy 7y agoI used to think this, and am still sympathetic to the motivation behind this. But I've since changed my mind when I saw the challenges faced in Big Data environments. Structured data is nice at first when you get into it. But the need for data prep stubbornly will not go away. If only because how your stakeholders want the data structured only sometimes aligns with the pre-defined structure. And when you get those numerous requests for a different structure, you're right back to data munging, and wrangling lines of text. In my limited experience, you're either taking the cognitive hit to learn how to grapple with line-oriented text wrangling, or taking the cognitive hit to learn how to grapple with line-oriented text wrangling AND on top of that a query language to fish out what you want from the structure first before passing it to the text wrangling. I'd sure like a general solution though, if you have a way out of this thicket.