4 ms·
part1 >> The problem with this approach is because figuring out where the files are requires knowledge of the tool inner workings, that can only be acquired fr
by aboytsov 14y ago
part1
>> The problem with this approach is because figuring out where the files are requires knowledge of the tool inner workings, that can only be acquired from reading the code or documentation
> I suppose this is true but it's really not an issue I have in practice. I run the pipeline and it produces (let's say) a .csv file as a result.
It's a good point and I, guess, I didn't mean it's a major issue. Just something which is, I believe, less than an ideal design, because it spreads (de-centralizes) information. For example, if you want to do something with the files outside of your workflow in some shell script, this shell script would contain a filename which a reader would have no idea how you came up with. Again, it's not something to obsess over, just an observation.
> There's really not a huge inconvenience in trying to find the output.
Even if so (highly doubtful in case of, as you say, 200 character filenames), this reasoning only applies to interactive sessions.
> Having the pipeline tool automatically name everything instead of me having to specify it is definitely a win in my case.
It's only true if you have to type less. I'm trying to make a case that you don't have to sacrifice clarity to achieve the same result. I'm trying to show you can win without losing.
> I suspect we're using these tools in very different contexts and that's why we feel differently about this.
That might be true, but we were also trying to come up with a universal tool. That is, we are willing to make sacrifices if not making them means severely limiting the scope of usage. But again, I am tying to show you don't even have to make sacrifices.
> It sounds like you need the output to be well defined (probably because there's some other automated process that then takes the files?)
Sometimes, yes; sometimes only for debugging; sometimes only for convenience. But more importantly, I'm arguing using filenames is just a better way to build the dependency graph regardless of whether you write them themselves or you use some identifiers that result in automatic filename generation. Remember I said Drake could easily do that? The core issue here is not filenames. It's what is the better (easier to read, less to type, easier to understand) way to define the dependency graph.
> You can specify the output file exactly with Bpipe, it's just not something you generally want to do.
Again, it's not the point. If you start specifying filenames exactly with Bpipe (I'm assuming you mean in commands themselves), you would just end up with a very strange beast: you'd have essentially define the dependency graph twice, once indirectly, and once directly. Or at least different dependencies in different ways. It seems like this would just be a total mess. But I'm trying to show even if you want to not care about filenames, Drake's approach is better.
> There's nothing wrong with either one - right tool for the job always wins!
My feeling so far was that it's not like a comparison of C and Python, but rather like a comparison of C and C++. There's absolutely nothing that you can't do in C++ better or at least as well as in C. Of course, I might be wrong, and that's why I would love to see an example workflow which I would then put in Drake and we'll be able to objectively compare.
> It just keeps appending the identifiers: will produce input.fix_names.fix_names.fix_names.csv. So there's no problem with file names stepping on each other, and it'll even be clear from the name that the file got processed 3 times.
First, I don't want to process the file 3 times - I didn't mean call the same method 3 times, I meant use the same code in different parts of the workflow. For example, you have a method to convert data from CSV to JSON, and you use it a dozen times all over the workflow.
Secondly, I think this is pretty bad. The way you described it, it makes filenames situational - i.e. depending on what part of the workflow they're in. Removing one fix_names from the chain could invalidate other fix_names's inputs and outputs, or worse - not invalidate the timestamps, but make such a huge mess, the user won't even know what hit him. Editing the workflow should not require such careful consideration for the tool's inner workings. And if you can afford to re-run the whole thing every time you add or delete the step, you're working on something very, very simple.
> One problem is you do end up with huge file names - by the time it gets though 10 stages it's not uncommon to have gigantic 200 character file names.
I apologize I didn't even realize the filenames carry all their creation history - I thought it was only the case with repeating names. I don't want to be harsh, but I think it's beyond bad. It means any change to the workflow can invalidate everything. This makes BPipe unusable for anything even remotely expensive. Please correct me if I'm wrong.
> Absolutely - you can get situations like this.
This is actually pretty common.
from(".xls", ".fix_names.csv", ".extract_evergreens.csv") {
exec "combine_stuff.py $input.xls $input1.csv $input2.csv"
}
I'm sorry, I tried but I didn't understand this code. Could you please elaborate? What do you mean "glob"? The way I see it, you may glob all you want, but there are just two ways to resolve this: use positional numbers or use some sort of identifiers. If you use positional numbers, it becomes unmanageable. And if you use identifiers, we're back where we started. It doesn't matter if they're filenames or not, what matters is that once you started using identifiers, you can generate the dependency graph yourself, from identifiers. In other words, you've arrived to Drake's model.
...continued in part2...