62 ms·
The myths of bioinformatics software
- coliveira 11y agoAn important point in this article is the vital distinction between academic software and "general purpose" software. The goal of research software is to prove or exemplify a point explained in one or more papers. It should do so in the most direct and economical way for the researcher(s). It makes no sense to create multi-platform, software-engineering friendly software for this kind of use. In the last few years we have seen an untold number of complains that research software is not robust, user-friendly, etc., and it entirely misses this important fact.
- angersock 11y agoThe problem is that the "direct and economical way" tends to basically enforce bit rot and lack a reproducibility. It also tends to screw whoever inherits the project.
- coliveira 11y agoReproducibility is an issue in all scientific fields, it is not something unique to computer science. Can someone easily reproduce experiments performed in a particle accelerator? What about an extremely complex chemical reaction? The answer to these problems is to have more people doing research in that area to make sure that the problem is well understood, instead of requiring that scientists slow down the research because it needs to first achieve some kind of "engineering reproduction" standard.
- bbgm 11y agoSpeaking for the chemical reaction - yes. That is exactly what should be possible. Some reactions are hard to reproduce and labs spend years trying to figure out how to do so. Without that you really don't have good science. The key is .. can someone read your paper and reproduce your experiment and the results. If not, it's not valid; it's just a paper.
- dalke 11y agoThat's not true. Some types of chemistry cannot (legally) be reproduced, but they are still good and valid science. For a trivial example, consider the many chemistry papers based on measuring fallout effects from nuclear testing. More prosaically, some things can be verified even when a chemical process cannot be reproduced. Protein crystallization is a tricky process with many irreducible results. Yet if you have a crystal, and can use the crystal to determine the protein structure, then you can verify that the crystal indeed contains that protein - even if you are unable to reproduce the chemical process used to create the crystal in the first place. Others can use the same sample to re-verify the resulting x-ray structure even though they also might be unable to reproduce the crystallization step.
- angersock 11y agoMy critique isn't aimed at computer science research--it applies equally well to other fields using computation to produce results. There are a lot of projects (and I won't name names, because I've got friends here who work on them) whose results cannot be reproduced, because they ran on some weird-ass combination of a weird libc, using a weird build system, running on weird hardware environments (especially common in HPC), with weird one-off data sets. They report the results using this setup, and then the setup is gone forever. The code is probably not kept correctly in source control, and generally the community just takes them at their word that the results they get are authentic. When you run millions of iterations on some crazy quantum solver, it's hard to spot if you're wrong. If others can't at least run the same result, science becomes a matter of anecdote.
- coliveira 11y agoAs hard to believe as it is for software engineers, science is not in the code, it is in the thought process that leads to the code and in the interpretation of the results. The code for an experiment can be lost forever, but if you did your job correctly you have described why, how, and where the experiment works. Future scientists can now "build on the shoulders" of this knowledge and use it to go further and validate the previous result or discredit if it is indeed incorrect. People complain that published papers are full of incorrect results, but they don't understand that this is just part of the process. Incorrect results will be forgotten because they will never be validated. Correct results will be validated and used (cited) by new scientists as they try discover even more complicated things.
- collyw 11y agoThings would work so much better if the software was written in a better way.
- stonemetal 11y agoThat seems rather short sighted. Science is often described as standing on the shoulders of giants, you can't do that if everyone has to start from scratch. Sure it doesn't have to follow NASA's coding guidelines but it also shouldn't be utter garbage that no one besides the authors can figure out.
- cossatot 11y agoThough you do have a point, I think it's really not binary like that; you've given end-members. There are a good number of scientists (myself included) who spend at least part of the time making applications for other researchers; it's best when the products are multi-platform, because there are a lot of scientists who only use one platform. These application can be anything from plotting and basic analysis tools to simulators of various sorts. They're more specialized than Evernote or something but are still designed to be installed and used by members of the community, because it's a waste of time for everyone to write their own finite element models (even if everyone knew how). When done right, not only are these tools enabling but they become very widely cited. For instance, in my field (geoscience) a tool called GMT (Generic Mapping Tools) has over 6000 citations, even though I'd estimate 70% of papers that have used it in some form don't cite it. This is an enormous number of citations in the geosciences; most very famous and influential papers have ~1-2000. Maybe it's different in CS; I have heard that much of the software is illustrating a new algorithm or something.
- deleted 11y ago[deleted]
- meeper16 11y agoAnd the most long standing myth, which first started after the Human Genome was sequenced by Francis Collins and Craig Ventor (now working on human longevity) over 14 years ago: Bioinformatics software will single-handedly be responsible for discovering billion dollar drug targets. This is in large part why most early bioinformatics companies failed - due to lack of deliver on this front along with jangled software approaches that were being moved into the commercial world from academia. The reality is that most bionformatics software relies on old formal methods and is not geared toward true highly innovative discovery and data interpretation. I do think however that we are entering a new age of bioinformatics and its associated data mining, interpreation, visualization and discovery tools which hopefully will push the bounderies of being less formal and more experimental. We need to make discoveries faster when it comes to Life Sciences. I think new approaches in bioinformatics/datamining/data science/visualization will have the greatest impact in the areas of extending human lifespan. This is what Craig Ventor, Google Calico Labs, SENS, GenoPharmix, Buck Institute are all working on now.
- kodisha 11y agoI also think it has to happen on the web (open or closed), instead of the desktop software. Also, d3 wont be able to pull it off. If we do a comparison, i think that d3 is MooTools, wee need to get to ES6/angular/react kind of libraries.
- collyw 11y agoWhat do you feel that ES6/Angular/React do that d3 won't? They are for building user interfaces. d3 is for charting. (Most bioinformaticians I know seem to use R).
- kodisha 11y agoIt was a comparison. Once MooTools was the library to use, now its dead.
- dalacv 11y agoWhy do you think it has to be web-based?
- rch 11y agoThe article mentions code quality and license issues, and one of my favorites (MEME) seems to suffer a bit from both. I believe the first aspect is simply the result of being developed in a sequential fashion by different contributors (which is reasonable given the environment). The main problem is that the license rules out using the software for 'commercial purposes' except under unspecified terms that would need to hashed out with the tech transfer office. I completely support the spirit of that construct, but it makes it difficult to advocate for in practice. At least in this case, GPL or LGPL would be a significant improvement.
- dalke 11y agoThe MEME commercial license is at http://techtransfer.universityofcalifornia.edu/NCD/Media/MEME%20suite%20site%20license%20template%2004June2015.pdf http://techtransfer.universityofcalifornia.edu/NCD/Media/MEM... . How is this "unspecified terms that would need to hashed out with the tech transfer office"? The license is US$2,500 for a license, which as Pachter correctly points out is a small cost for most companies. Also, while it would be a significant improvement to you, would it be a significant improvement to the science? For example, I can't speak to MEME but I know of a couple other projects where the software is at low/no cost to academics and has a license fee for commercial use. This money is used to fund future development, which gives a funding source that is independent of grant funding. Pachter also points out this possibility. It can be frustrating if some people cannot use a package due to license terms or costs. But it can also be frustrating to use a package where no one is available to answer questions or fix bugs - which is something that funding can address.
- bbgm 11y agoIf you are bootstrapped it is a huge barrier. And Meme is not that bad. There are others that are far worse. I once tried using Modeller purely for hobby projects (science was not even my day job) but was denied. That's a problem.
- dalke 11y agoNot all business models are economically viable, and others are not obligated to make it easy to support your choice of a bootstrap business model. My experience is that people who get a piece of software, even if for free, often want some support. If you don't give them support, some will complain. A company may decide that it's easier to deal with complaints about the lack of a free version for hobbyists than to deal with complaints about a user not getting sufficient (unpaid) support.
- cjbprime 11y agoAs https://twitter.com/madprime/status/619503684838387716 https://twitter.com/madprime/status/619503684838387716 points out, the argument that you're cheating the US Government out of public money by releasing without a non-commercial clause is bizarre -- everything the US Government releases is required by law to be released into the public domain.
- maaku 11y agoExcept that's not true? Most things release "by the government" had a contractor involved somehow, and government contractors have a different set of rules.
- i000 11y agoAre you suggesting scientists are "goverment contractors"? Never thought of myself as one.
- meeper16 11y agoWith the big exception of what is funded by the DOD/DOE, DARPA etc and run through national labs like Berkeley, Livermore, Sandia... this is why they have commercial licensing and tech transfer operations in place.
- GFK_of_xmaspast 11y agoThere's a huge difference between 'government employee' and 'accepts federal grants'.
- jerven 11y agoYour data is more important than your code. Is the often neglected fact in bioinformatics. Whatever you do document your file formats.
- roel_v 11y agoMuch of this is equally applicable to other fields, but I very much disagree with point 2. Every lab should have one or more programmers, people who are professionals at writing software, and who guide researchers in their software development. For both efficiency and accuracy reasons. But of course a software developer at a university is, at best, a 'lab assistant', but more likely regarded to be on the same level as the janitor (both in respect and in pay). With the result being thousands upon thousands of shitty programs you wouldn't wish work on to your worst enemy. But hey, cool with me, I carved out a consulting niche in cleaning up such messes in exactly that environment. But man could a lot of money be saved, and a lot of much better work be done, if only researchers (and the hierarchy above them) would recognize that software development is both critical to pretty much any research today, as well as something they cannot just pick up on the side.
- wlievens 11y ago> But hey, cool with me, I carved out a consulting niche in cleaning up such messes in exactly that environment. Sounds like an interesting story there. Anything you can share about it?
- roel_v 11y agoNot so interesting really, just that funding is structured in such a way that there are ways to get funded to convert research prototypes into production-quality software, and as long as you have researchers willing to include it in their proposals, domain knowledge to have credibility and hold your ground in your consortium, and a network within which you're known for this; then you can make a living as an independent researcher/software consultant. I'll add to this that I stumbled into some peculiar, non-replicable circumstances which let me bypass the 'only eat ramen for several years' phase that is the rite of passage in the academic world. The opportunity cost of that would have made it non-rational for me to do this. So I don't have actionable career advice on this line of work, I'm afraid.
- collyw 11y agoI am a software engineer for a sequencing centre and I completely agree. Scientists don't value proper software development and think that writing one off scripts is the same thing.
- shiggerino 11y agoInsisting bioinformatics software be non-free is pretty rules out any possibility anyone is going to build on your code. If this is the case, that's regrettable, but why seal the fate? If they are afraid of companies using and abusing the software, just put it under the GPL and they will at least have repay the favour to the users and the community.
- dalke 11y agoThe essay gave an example of non-free project that others have built upon: > One of the most widely used software suites in bioinformatics (if not the most widely used) is the UCSC genome browser and its associated tools. The software is not free, in that even though it is free for academic, non-profit and personal use, it is sold commercially. ... As far as development of the software, it has almost certainly been hacked/modified/developed by many academics and companies since its initial release (e.g. even within my own group). Therefore, by demonstration, using a non-free license does not "seal the fate" and rule out others from building on your code. Also, GPL does not require anyone to "repay the favor." There's no requirement to distribute modifications upstream or to "the community."
- shiggerino 11y ago>There's no requirement to distribute modifications upstream or to "the community." No, that would obviously be a onerous requirement. The GPL strikes a reasonable balance between the individual user and the community of users. I'm just saying the customers should be allowed to get the derived software on the same generous terms as the company received the original on. The customers are obviously not required to redistribute anything, but it encourages good behaviour.
- dalke 11y agoIf it's "obviously [an] onerous requirement" then what does "they will at least have repay the favour to the users and the community" mean? You clarify that "customers should be allowed to get the derived software on the same generous terms as the company received the original on", but that assumes that the companies have customers. Very few companies that use academically produced bioinformatics software have downstream customers of that software. For the vast majority of companies that only use the software in-house, which is likely 99+% of all companies, how does using the GPL or any other free software license lead to the company repaying the favor? What's wrong for asking for payment in cash instead of other more nebulous contributions?
- bmir-alum-007 11y agoDisclaimer: I used to work at a Stanford bioinformatics shop. There's a clear need of AWS-like features for bio/biomedical informatics specifically enabling sharing, security, reuse and anonymization of data (PHI), libraries (like R's bioconductor) and infrastructure (IaaS/PaaS/SaaS). The issue is that some labs archive still archive their data on actual hard drives (USB and bare drives), making their data much less useful than somewhere readily available and sharable. I think it's a huge (billion+) opportunity where the right execution would need loads of smart, consultingish customer service reps (huge overhead costs) to help researchers with coding, sysadmining and bio to some degree. Basically, a full-service (with self-service, a-la carte features) hosting company for bio / medical. This space is only going to grow deeper and wider as more is discovered and confirmed about each gene, protein, pathway and each accompanying expansion in nosology. This sort of research knowledge is vital and unlikely to shrink. The main issues are that it would be a cash-intensive and undefensible business model because it requires paying lots of consultant/scientist brains and anyone can copy the model.
- x0x0 11y agoand you're competing with dirt-cheap grad students / student ras / post-docs most labs are just too cheap (I wrote custom stats software for image / FRET analysis and my lab was certainly too cheap.)
- infinite8s 11y agoThe main issue is this expertise is too expensive to provide at the lab level - probably needs to be provided at the institutional level.
- macarthy12 11y ago> There's a clear need of AWS-like features for bio/biomedical > informatics specifically enabling sharing, > security, reuse and anonymization of data (PHI), libraries > (like R's bioconductor) and infrastructure (IaaS/PaaS/SaaS). I tried to build a startup like this, basically a Heroku for bioinformatics, with a bunch of experienced biologist / genome folks. They just didn't get it and on my part, I guess I couldn't sell it to them. Part of it was snobbery, and institutionalized thinking. It was a big disappointment. Some one will do it, but until then it will crappy bioperl scripts, with no version control etc.
- arca_vorago 11y agoThe first opportunity to comment on bioinformatics since my non-compete/nda is over! This seems like a very sloppily put together list of myths, but I'll bite anyway. 1. Not true, but I think that's largely because much of the software used is closed so the FOSS community is largely anemic in the bio world. For the tools that are FOSS or BSD, I saw plenty of contributions, but the other thing to keep in mind is that it's not just about the programming. You have to have a certain level of understanding of the application domain to program a solution for it properly, and there are very few of these people around. I predict a huge uptick in demand and salaries for bioprogrammers. 2. Is true. You need your own people on salary to program for your needs. I was the sysadmin part of a phd, sysadmin, programmer team and we were doing stuff that no-one else was going to do for us. You need to have your own programmer, and a good sysadmin, full stop. 3. Is also true. Picking the right license is important because many labs are pretty tight on cash flow. Sure, they probably have millions going through them a month, but operating costs are super high and margins are lower than you may think. It was during my time in the genetics lab that I fully realized why FOSS was so important, and I think it's the future. (with a few key proprietary exceptions that no FOSS has matched yet, (think Elmer vs Comsol)) 4. Using a FOSS license makes this a moot point to address. Use GPLv3 code people, stop using BSD! 5-9: not worth addressing. Anyway, my overall view of the field is this: with sequencing getting cheaper, the problem is in managing the levels of data being generated (sysadmin issue) and in interpreting the data for meaningful results (programmer/phd issue). Personally, I think that machine learning is going to be the right breakthrough to follow and apply to bio, and once we do that I expect it to take off to crazy levels. I'm talking sequencer in every doctors office, and artificial genetic manipulation becoming much easier and with more accurate predictions. Also, the other thing everyone underestimates is the microbiome as an entity. You are more the bacteria that lives in you than you are you. Of course, I struggle to understand the science sometimes, I'm just a sysadmin, so take what I say with a grain of salt.
- dragonwriter 11y agoYou may be confusing "FOSS" (Free/Open Source Software) with copy BSD is F/OSS, it is not copyleft. GPLv3 is both F/OSS and copyleft.
- danieltillett 11y agoAs someone how actually makes a living selling bioinformatics software, the problem is mainly due to how scientist view software. The code you write is seen the same way lab books are - basically raw data. Nobody publishes their lab books and all too often software is thought of as just an electronic lab book. It would be great if this changed, but it needs a change in how scientist look at software.