4 ms·
I've considered using Nextflow for bioinformatics pipelines but have yet to take the plunge. At work, I develop a proteomics pipeline that is composed of huey¹
by radus 3y ago
I've considered using Nextflow for bioinformatics pipelines but have yet to take the plunge.
At work, I develop a proteomics pipeline that is composed of huey¹ tasks (Python library; simple alternative to Celery) which either use subprocess to call out to some external tool, or are just pure python. It runs in a worker container which is managed by Docker swarm, and all containers pull jobs from redis. For our scale, it works great. However, I don't have control over the resource utilization of individual steps, and in the past I've had issues with the pipeline blocking as a result of how I was chaining tasks together. I think something like Nextflow would remove these limitations, but one thing I think I would miss is the ability to debug individual pipeline steps locally with an interactive debugger. As far as I can tell, Nextflow has logging/tracing facilities but nothing quite like an interactive debugger. I'd be happy to be told I'm wrong, or even that I'm doing it wrong.
Other reasons I'd like to start using Nextflow:
- my homebrew pipeline would be easier to setup/share
- there are some efforts in the proteomics community to develop Nextflow pipelines (eg. QuantMS²). I think it would to have a shared language to express pipelines, and it would make benchmarking simpler.
___
¹ https://github.com/coleifer/huey/ https://github.com/coleifer/huey/
² https://docs.quantms.org/en/latest/ https://docs.quantms.org/en/latest/
- nonrepeating 3y agoThe closest I’ve gotten to local debugging is having the Python scripts that are launched by NextFlow steps connect to a remote debugger process (“remote” but running on the same workstation). PyCharm makes this fairly painless to orchestrate. I’ve never been able to debug thr Groovy script in a Nextflow pipeline itself; I think you’d need a debug build of the nextflow executable for that.
- a_bonobo 3y agoAs someone who dips into nextflow from time to time, I'd strongly suggest by developing your pipeline based on an existing nf-core pipeline or the nf-core templates. nf-core comes with a bunch of nicer defaults like profiles for SLURM, Singularity, Docker that help you abstract some of the headaches away, plus you could get lucky and can just glue some of their modules together.
- geoffjentry 3y agoI don’t know. I started this route and then quickly switched to only dipping into nf-core when they had actual prior art. The interplay of nf and groovy (how I wish they hadn’t used groovy!) can be mind bending but if you’re writing your own thkng you have a different optimization model than nf-core that is trying to be one size fits all
- biophysboy 3y agoYou can debug snakemake with pdb. It also has actual dry-runs to test the dag before actually running anything (with nextflow you have to test run with “stubs”)
- getoffmycase 3y agoFrom one proteomics person to another, what tools are you using? I can see needing snakemake for something like proteogenomics (our lab published a tool in that area) or DIA cause that pipeline can get a little complex. But for run of the mill stuff, as long as I have CLIs, I don’t really find myself needing anything beyond a basic batch file. Before you wonder why I don’t know that, I do top-down software development primarily and also maintain and upgrade my lab’s search engine, MetaMorpheus.
- radus 3y agoThe DDA pipeline goes like this: ThermoRawFileParser -> Comet -> [a bunch of OpenMS tools] -> Percolator -> [custom quantification stuff]. The data is mostly derived from chemoproteomics experiments where you have isotopically labeled control and compound treated samples that are enriched with some probe. As a result, we work a lot with ratios and have to differentiate the scenario where your compound completely blocks probe labeling/enrichment or it's just stochastically missing due to the nature of DDA. For TMT it's pretty similar. I'm working on DIA as well though it turns out there's still quite a few challenges there for our particular use case. To answer your broader question about the general need for some structured pipeline or workflow orchestration.. That comes down to volume of data (we do screens as well as one-off studies) and a desire to reduce human involvement as much as possible. So the goal is to have raw files be immediately picked up, processed, and loaded into on internal application where it can be queried and interesting data can be highlighted. During my PhD, this was also a goal of mine (and I have at least two github repos where I got close) but it was definitely less of a priority since actually doing experiments and downstream analysis was the limiting factor. PS: if you want to talk off-HN, I should be your latest stargazer
- getoffmycase 3y agoI’d be happy to chat offline. Your current work and my current work I think share a lot of similarity. My GitHub is avcarr2.