3 ms·
That's a big part of it. And I say this as someone who played a heavy role in the development of one of the named DSLs. Yes we had that exact intention and exac
by geoffjentry 5y ago
That's a big part of it. And I say this as someone who played a heavy role in the development of one of the named DSLs. Yes we had that exact intention and exact unintended outcome.
There's another part of it as well. Bioinformatics workflows tend to be just enough different from more standard workflows that there is friction using off the shelf tooling.
For one, the DAG nodes are often mostly/fully represented by command line tools expecting a POSIX style file system and making assumptions/asserting opinions on where files live, to where it writes outputs, etc. Bioinformatics workflow orchestrators can understand this and provide optimizations in the DSL to express how to manage the file movement. In contrast, I find many people with a more standard data/biz workflow mentality think of DAG nodes as being queries, running blocks of code, etc.
Lifecycles vary as well. The institutions who have these workflows will either be pure research or some degree of research/production hybrid. There's an advantage to being able to use the same software on both side of that hybrid. When your research workflow is ready to be blessed to production, having to translate it into a different system can be expensive in terms of both time and bugs.
Another aspect is that these workflows, until not too many years ago, were often run on HPC compute clusters using software like SGE, SLURM, PBS, etc. These DSLs can provide optimizations for tweaking parameters in a way that more mainstream tools do not.