3 ms·
Fellow bioinformatician here. I find it amusing to see all these projects reinventing pipes and makefiles much more verbosely and without language independence.
by xaa 12y ago
Fellow bioinformatician here. I find it amusing to see all these projects reinventing pipes and makefiles much more verbosely and without language independence.
The project in the OP, it goes without saying, will never be used by more than a dozen wet-lab biologists, and that's being generous. Possibly it may have some use in automation.
Wet-lab biologists have many other things to worry about, and do not have the incentive or technical knowledge to record all the parameters of their experiments in a complex computer system. If the incentive existed, this data would be sent in Excel spreadsheets to NCBI, and NCBI would turn it into some half-assed file format.
- samuell 12y agoI'm sure makefiles go a long way (many people solve their needs with them), but there are also a lot of use cases where they start to fall short, AFAIS. Think e.g. a cross-validation set up, combined with a parameter-grid search (to find an optimal parameter combination, for, say, building a support vector machine model), where certain re-usable workflow components (training and prediction) are run for each parameter combination, for each fold in the cross validation ... We are doing that kind of stuff, and couldn't really imagine any sane way to implement that in make ... which made us go with Spotify's luigi, as documented at https://medium.com/@saml/loosely-coupled-tasks-in-luigi-workflows-6840d32e2824 https://medium.com/@saml/loosely-coupled-tasks-in-luigi-work... I still could imagine making this much easier using a light-weight system such as the mentioned "blow".
- xaa 12y agoYou are right that makefiles alone would be a bad fit for that situation. My approach would probably be to wrap the parameterizable parts in a script accepting arguments and run the combinations with GNU parallel. Possibly controlling the overall execution and dependencies in a makefile. It's less "clean", because you have to keep track of metadata like parameters in the output filename or similar. But the advantage is a huge increase in flexibility. We tried Celery for awhile as a job manager which looks like it has similar capabilities to Luigi. It was slower, caused us to write a lot of ugly "bash-in-python" when calling non-Python programs, and broke the UNIX philosophy of having independent programs doing one thing, making it harder to quickly test new combinations of components without writing a lot of Python code. It also depends on your dataset size. We do a lot of machine learning on datasets that won't fit in RAM, which is perfect for the pipe/streaming model.