4 ms·
> I'm sorry I don't have time to answer in full. We're not getting anywhere. Just give me goddamn examples! :) Please! Examples! > I can see this is really re
by aboytsov 14y ago
> I'm sorry I don't have time to answer in full.
We're not getting anywhere. Just give me goddamn examples! :) Please! Examples!
> I can see this is really really hard to grok if you're basing everything on the idea of a DAG, and so many tools are that it's very natural to think you couldn't do it any other way.
There is no other way. BPipe is based on the idea of a DAG. You just don't see it.
> In Bpipe the user declares the pipeline order explicitly.
And this is a big mistake. The reason is simple - explicit order is very hard to manage once you have multiple inputs and outputs, and as a consequence, complicated (instead of linear) dependency relationships.
What you don't seem to realize, is that by "declaring the pipeline order explicitly" you create a dependency graph. It's a part of your workflow definition. Your workflow contains the full definition of the dependency graph. Even if it didn't, you would still use it. There is no other way.
This is what I meant when I said - you create your dependency graph in "run". And this is a bad idea.
> dependencies arise as actual commands are executed.
What does it mean exactly? That the first command will somehow tell Bpipe what to run next? If not, then I don't understand this statement at all.
> How is that if it doesn't know about the dependency graph?! Well, it does it "just in time".
It does not matter if you calculate the dependency graph before you run the first command, or as you run the commands. It makes absolutely no difference. The only difference is whether it is computable or not. If you say it's not computable until run-time, please elaborate on that.
> So in this way Bpipe handles dependencies for you.
So far I see that this is very standard and doesn't differ in any way from what Drake or any other tool does. The only thing that differs, and I am repeating myself, is how you define your dependency graph - through input and outputs, or in "run". So far it seems that "run" is quite unfortunate. But please give me examples.
> So in this way Bpipe handles dependencies for you. What it does not do is figure out which order to execute things in. It does them in exactly the order you tell it.
This is a meaningless statement. Drake also executes steps in the order you tell it. The only difference is how you tell it. In Drake, you tell it through specifying a list of steps each step depends on individually (once again, it doesn't matter that filenames are used for that - Drake also supports tags, or it could be some other identifiers). In Bpipe, you tell it in "run", collectively and sequentially. Drake's way supports the whole variety of graphs, while Bpipe's way - only a very limited subset. And for this limited subset, Drake can give you (I think) a syntax just as good if not better than Bpipe's. If you don't quite understand what I'm talking about, give me an example, and I will demonstrate.
> I actually want to control the order of things sometimes.
This is fine, the only question is how. You say Bpipe's way is convenient. I say give me an example and I'll show you that Drake's way is not any less convenient. I'm sorry to keep repeating myself, I thought I stressed the importance of examples quite a bit in my previous email and I want to stress it again. Examples, please!
> I want to be able to tell it "do this first, then that, then the next thing" regardless of dependencies.
This statement is self-contradictory. You don't seem to realize that by telling it "do this first, then that" you are defining dependencies. It's fine, and it's OK, and it can be convenient, but you can't say regardless of them.
Again - give me examples! Our conversation is becoming useless without examples.
You did not, but I'll just grab whatever you threw my way:
fix_names = {
exec "sed 's/Neverbrown/Evergreen/g' $input > $output"
}
extract_evergreen = ...
run { fix_names + extract_evergreen }
$ bpipe run pipeline.groovy input.csv
Drake can support this perfectly:
_ <- $[in]
exec "sed 's/Neverbrown/Evergreen/g' $INPUT > $OUTPUT"
$[out] < _
........
$ drake -v out=pipeline.groovy,in=input.csv
Isn't that much nicer? What disadvantages you can see?
Tell me what is it that you would like to do with this script, and I'll tell you a better way to do it in Drake. Is it multiple versions of run that you want to have? Easy. Are you concerned about inserting a step in the middle? Trivial. Tell me why Drake's code is worse, and I'll listen. So far it seems like it's better because it's shorter and more flexible at the same time.
> Having the tool think this stuff up by itself can save you a bit of time but it can lose you a lot because you don't have the ability to really control what's going on.
What exactly are you losing?
I am sorry if I sound irritated. I am. I've just been begging for examples, and you keep talking in abstract, and it would be fine, but you're making a lot of mistakes. So, instead of looking at concrete things that would make my point apparent to you (or the opposite, prove that I'm wrong), I keep pointing to flaws in your reasoning, which frankly, is irrelevant. One picture is worth a thousand words.
I really want your feedback. But please give me examples.
- zmmmmm 14y ago> There is no other way. BPipe is based on the idea of a DAG. You just don't see it. So if you think Bpipe uses a DAG, then I wonder how you would think it deals with: run { fix_names + fix_names + fix_names } In terms of the pipeline stages that run this is cyclic, so it cannot be a DAG. On the other hand the files created do usually form a DAG dependency relationship, but even there, in the most general case, it's not at all impossible in an imperative pipeline to read a file in and write the same file out again in modified form (or more likely, to modify it in place), so the file depends on itself - another non-DAG relationship. I'm sure you'll object to this in a purist sense, and tell me it is a horribly broken idea, but as a practising bioinformatician, when I have a 10TB file and modifying it in place will save me hours and huge amounts of space, I'm much more interested in getting my job done than being pure about things. I think you're right that we're at diminishing returns here, and I'm sorry I've frustrated you. We're trying to bite off more than we can chew in a forum like this. I wish you all the best with Drake and I'll definitely check it out down the track (when it supports parallelism, since that's too important to me right now). For now, though, I don't intend to read / respond to any more replies in this thread.
- aboytsov 14y agoThis is not a cyclic dependency graph!!! This is a syntax for copying vertices, nothing else. It creates a DAG of three vertices and two edges, but uses only one step definition to do so. It automatically replicates the step definition as needed. It could be extremely easy to reproduce in Drake: fix_names() ... _ <- $[in] [method:fix_names] _ <- _ [method:fix_names] $[out] <- _ [method:fix_names] Is there any difference between Bpipe's version and Drake's version that I am failing to see? > I'm much more interested in getting my job done than being pure about things. It's funny coming from someone who I have been BEGGING for examples but getting abstract philosophical reasoning in return. I repeat. Give me an example. So far you haven't given me one example of what Bpipe can do that Drake couldn't do in the same way or better, and yet you continue claiming philosophical differences. If we concentrate on examples and discuss how they would work, whether there are differences, and what these differences are, I guarantee you, we'll make progress. But then again, I'm repeating myself. Artem.