8 ms·
Recently, I was helping out a friend in analyzing some RNA samples for her work. These samples are huge - like nearly a gigabyte of data. There was this tool wh
by kumarharsh 8y ago
Recently, I was helping out a friend in analyzing some RNA samples for her work. These samples are huge - like nearly a gigabyte of data. There was this tool which was recommended for the job - mirexpress. It was a small job, perhaps 10 minutes worth of effort. To make my work easier, I provisioned a beefy (and costly) machine on Azure to do the job, took a quick look at the clock (it said 11 PM), ran the tool, and relaxed. The tool crashed while reading the file.
In an attempt to fix the bug, I opened mirexpress's code. And all my confidence in my programming ability vanished when I saw its innards. I understand that the code may have been written by scientists who had no experience in programming, but I have never been so utterly _disoriented_ by bad code. Anyways, after hacking away at the mess for about 3-4 hours, I realized that this was a fool's errand and thought I'll just phone it in the next day saying I couldn't do it. I went to sleep thinking that it was already late and I'd get late for work the next day.
- 5 minutes later -
I woke up with a start, recalling this nifty tool called awk. I had last used it maybe 3 years ago, and before that only in college. But I could see how awk could do some of the things which mirexpress was claiming to do. So I fire up my computer, write an awk script - 2 lines only! TWO FUCKING LINES! And it runs like a charm - eats away at megabytes of sample data and gives me results I can show. So then like any rational person, I spent the remaining hours re-discovering awk and forgot to sleep. Pissed away the whole next day (and some part of the day after that too!) :-D
It's really fascinating that this nifty little tools invented DECADES ago are still going strong, and there's been no _evolutionary_ leap in areas where tools like awk/grep/sed excel at.
- omginternets 8y ago>but I have never been so utterly _disoriented_ by bad code One of these days, when the mood strikes, I'll upload the matlab script I inherited from my Ph.D supervisor. Scientific programming is, by and large, an abysmal state of affairs.
- Waterluvian 8y agoAt one job I had we got "help" from a local college professor on how to design our database. A lot of what he said made sense. And we were super green and databases were not our expertise or focus. Short on time, high on cash, we asked if he'd implement it for us. I will never forget my gut wrenching feeling when I opened up one of the main provisioning scripts (in Python) to read the first two lines: True = 0 False = -1 The rest of it was an unmaintainable disaster that may have worked. I learned so much that summer, having been forced to learn it all myself.
- djtriptych 8y agoI literally almost choked on my drink reading this.
- felipemnoa 8y ago>>I will never forget my gut wrenching feeling when I opened up one of the main provisioning scripts (in Python) to read the first two lines: True = 0 False = -1 << Could you please explain what is so egregious about those two lines for those people that may not see it? Asking for a friend.
- kpcyrd 8y ago>>> True = 0 >>> False = -1 >>> if False: print('oh no') False is now truthy and True is falsey. Note this only works in python2.
- wingerlang 8y agoTrue is normally 1, while False is 0. This seems like specific variable declarations as well, which you generally don't need as they are built into all languages e.g. "variable = false;" However maybe the professor needed 'true/false' to represent something to be used in his formulas so this made sense to him. But who knows.
- Buge 8y agoIn most programming languages, including python, true is represented by 1 and false is represented by 0. This script is changing it to be backwards, and python is flexible enough to allow that. But I would expect this would break most python libraries, because it's such a core language feature that's being completely restructured. https://en.wikipedia.org/wiki/Boolean_data_type https://en.wikipedia.org/wiki/Boolean_data_type
- deleted 8y ago[deleted]
- poxrud 8y agoWhen doing comparison checking, in most programming languages false is 0 and true is any other value.
- fwip 8y agoStandard unix tools are so incredibly useful in bioinformatics. Almost all of our data formats are, or can be canonically transformed into, plain-text. Learning the basics of awk, sed, etc is almost a requirement.
- skosuri 8y agoSort, cut and uniq are also amazingly useful.
- cup-of-tea 8y agoDon't forget paste. You need that to make your table from your column oriented data store.
- fatboy93 8y agoDon't forget join, tr and column. Makes formatting a breeze. Most of my time is spent on the terminal and without these I'd shoot myself.
- AceJohnny2 8y agoNice story! I have a friend who worked in bioinformatics, but he was good at programming and went to better-paying pastures. I wonder if the problem is twofold: 1) a lack of education compounded by 2) the rapid evolution of computer systems. Unix is a rare beast in that not just its philosophy but even its components survive to this day and remain relevant[1]. People in unrelated fields rebuild tools that could just as well be assembled using Unix's basic components, but they're just not aware of them. And why would they look for these antiquated tools? They've been trained to reasonably expect old tools to have been replaced by newer, better, more featureful ones. [1] AWK was created before I was born, and I'm among the more senior engineers in my 20+ team.
- kumarharsh 8y agoI think it's just that they are good at biology, not computers. I mean if they entrusted ME to tell two amino acids apart by looking at their behaviour, I'd take a month! From what I've seen, amid all the research, failed experiments, constant fight for funding (in some cases) and "life" in general, people have much less time to learn these critical things like programming / tools so useful to them. They will do it at some point, but probably not till their life depended on it. On the other hand, if life depended on me recognizing those amino acids, we'd all be :-)
- ENGNR 8y agoI think part of it too is that it's exploratory. Sometimes I use javascript to write music and as creative twists and turns happen the code folds in on itself in ugly ways that don't happen when solving a clearer problem Compounded by the fact that in science once you have the result the code is probably just an artifact so there's no real reason to refactor it
- jecxjo 8y agoWhen you start looking at the code of the tools, and how older systems were design, it's really is a testament to good engineering practices because it all scales so well. Small, simple, "one task well" utilities that process data in streams means the megabyte files of the 80's and 90's are today's terabytes. Sure they require some work to get things scripted correctly but it really is pretty amazing what coreutils and a shell script can do.
- molotovbliss 8y agoThere Unix philosophy. https://en.wikipedia.org/wiki/Unix_philosophy https://en.wikipedia.org/wiki/Unix_philosophy It's a shame to see so much bloated, over abstracted development these days with more emphasis on delivery than doing one thing well. https://en.wikipedia.org/wiki/Demoscene https://en.wikipedia.org/wiki/Demoscene After seeing what the demo scene is able to pull off in very little code makes me hope that the art of fundamentals & single responsibility principals like in Unix/Linux resonates with younger developers the ability of truly understanding a problem scope before implementing a solution. Above all else have fun with it, be curious on even a low level how your program functions at an os & hardware levels. Things like what is the von Neumann bottleneck? Or even those building boolean logic gates in minecraft is inspiring to want to know "how does it work? And why?"
- bleachedsleet 8y agoYou see this in malware development a lot as well, the search for smaller and more efficient executables that use a minimum of code to function.
- molteanu 8y agoAlso a nice read: The art of writing Linux utilities (Developing small, useful command-line tools) [0] [0] http://people.fas.harvard.edu/~lib113/reference/unix/writingtools.html http://people.fas.harvard.edu/~lib113/reference/unix/writing...
- vram22 8y ago
- rspeer 8y agoThe one thing I find unfortunate about how people learn awk is the emphasis on making it terse and unreadable, especially trying to turn everything into a shell one-liner. Don't write one-liners if they're not that simple. Write ten-liners! Add comments! Commit them to version control! Awk can be a nice, readable little language if you're not trying to win at code golf all the time.
- chrisfinazzo 8y agobwk addressed this in a video at U Nottingham a few years ago. In a nutshell, he's basically arguing that they designed it to be terse to avoid giving people an excuse to write 1000's of lines for no reason. Less code = better (at least to his mind).
- rspeer 8y agoI can't imagine what one would do with thousands of lines of awk, so that seems a little extreme. Awk is a pretty terse language, and that gives you plenty of room to put in whitespace and comments and still have something short and sweet. I think that the idea of writing code to be nicely readable by your collaborators, or your future self, might not have been around in the early days of UNIX. ("But it's just a one-off thing I'm never going to do again, why should I save it to a file?", you may ask. Is what you do important? Then you'll probably have to do something like it again.)
- bsg75 8y agohttps://github.com/lh3/bioawk https://github.com/lh3/bioawk
- kumarharsh 8y agoWow, thanks!
- pbhjpbhj 8y agoack-grep has been kinda revolutionary for some of my uses.
- fatboy93 8y agoI fucking love awk and sed so much that to my chagrin I haven't learnt to program properly. I tend to one line most of my things in bash and tend to write a script if I have a multiple use case. Such is the folly in bioinformatics. Most of the stuff that we use tend to be one off. :/
- samuell 8y agoI counted to at least 10 awk components in a recent bioinformatics/cheminformatics workflow we wrote (Search for 'awk' in these lines: https://github.com/pharmbio/ptp-project/blob/aa91f1/exp/20180426-wo-drugbank/wo_drugbank_wf.go#L177-L329 https://github.com/pharmbio/ptp-project/blob/aa91f1/exp/2018... ). For the curious, the workflow builds ML models to predict binding to unwanted proteins for new drug compounds, based on drug molecule-protein target interaction data, and is implemented with out Go-based SciPipe workflow library [0]. The rawdata is a 18GB tsv file (ExcapeDB [1]), and AWK really really shines for this. I sometimes wish we had an SQL-like language for expressing all these computations, because we ended up doing some pretty complex join-like filtering and stuff. But from my tests with SQLite, I would not get a query answer until after like ~5 minutes with SQLite, so I gave up. With AWK, I can do a "head -n 10" on the result of the AWK operation (or on the input data file), and immediately verify that it works as expected. This seems to be one of the strongest points of the unix tool philosophy: Ability to check partial results in no time, and so iterate quicker on developing long "pipelines" of chained commands. It also fits perfectly as a way to build scipipe workflow components, since having the component code in a separate language (AWK), rather than as inline Go, which is also possible in SciPipe, it is super-easy to send this component of to the HPC resource manager (given that we have a shared parallel filesystem): Just prepend the SLURM [2] salloc command with its parameter, before the command. [0] http://scipipe.org http://scipipe.org [1] https://zenodo.org/record/173258 https://zenodo.org/record/173258 [2] https://slurm.schedmd.com/ https://slurm.schedmd.com/
- nicoburns 8y agoSounds like xsv might be a nice complement to your toolkit https://github.com/BurntSushi/xsv https://github.com/BurntSushi/xsv
- samuell 8y agoLooks very interesting, thanks!
- cat199 8y ago> an SQL-like language for expressing all these computations you might be interested in datajoint: https://github.com/datajoint/ https://github.com/datajoint/ https://tutorials.datajoint.io/ https://tutorials.datajoint.io/
- jhbadger 8y agoAlthough standard awk has the advantage of being installed on all POSIX systems, for bioinformatics there's a nifty bioawk (https://github.com/lh3/bioawk https://github.com/lh3/bioawk) that extends awk so that records aren't necessarily lines but fasta/fastq records as well. Pretty neat for oneliners.
- wizardofmysore 8y agoag was a development over grep.
- hx2a 8y ago> So then like any rational person, I spent the remaining hours re-discovering awk and forgot to sleep. Pissed away the whole next day (and some part of the day after that too!) :-D Haha! I've been there! Awk is amazing and when paired with sed it exponentially increases what one can achieve on the command line.
- ratsmack 8y agoI used mawk for many years and wrote numerous >1000 line programs for manufacturing automation. It is just about the fastest interpreter I have ever encountered, processing multi-megabyte files in just seconds.