3 ms·
Please see my response to Make comparison: http://news.ycombinator.com/item?id=5111527 http://news.ycombinator.com/item?id=5111527 I suspect most of the point
by aboytsov 14y ago
Please see my response to Make comparison:
http://news.ycombinator.com/item?id=5111527 http://news.ycombinator.com/item?id=5111527
I suspect most of the points I made would be applicable to redo as well, if not more so. Trivial things don't require Drake. Heck, they often times don't require Make as well - just put it in a linear shell script if the steps are not too expensive. It's when things are getting complicated you need something like Drake.
- moonboots 14y agoRedo lacks features baked into Drake, especially the Hadoop integration, but I believe it would be easier to incorporate custom functionality into redo versus hacking Make or writing a custom build system. I haven't used Drake, so I would be interested in a small but complicated Drake script which tackles an intractable problem in Make. I don't claim redo can provide a cleaner solution than a purpose-built system, but I think it will be unexpectedly simple.
- aboytsov 14y agoThe most crucial thing that Make lacks is multiple outputs and precise control over execution. When you're debugging/developing a large and expensive workflow, you absolutely must have the ability to say things like: - run only this step, I'm debugging it - I've changed implementation of this step, re-build it and everything that depends on it - build everything except this branch, it's expensive and I don't need to rebuild it that often (example: model training) Other examples of intractable problems in Make would be timestamped dependency resolution between local and HDFS files. If Make can't look at HDFS, it can't say if the step needs to be built or not. I don't think you can fix it with external commands. But generally, search for intractable problems is a futile one. Remember, everything you can code in Java, you can code in a Turing machine. :)
- lars512 14y agoMake can certainly generate multiple outputs, and can trivially be coerced to redo any step you like. Provided you add your code as a dependency in the analysis, then it will happily redo only what's changed, giving you nice tight iterations. I think it's real limitations are with multi-machine setups, as in the HDFS problem you're mentioning. Then you need a new tool.
- aboytsov 14y agoSorry, I might be very ignorant of make - could you please give me a command to re-build a particular target and everything that depends on it?
- lars512 14y agoSo make's default behaviour "make somefile.csv" is to build the whole tree of dependencies. To force rebuild of everything, run "make -B somefile.csv". It then assumes everything is out of date. To force rebuild of one step, just delete its output or run "touch" on one of its dependencies before running make. Then that step will get redone. I like to have generated data in a separate folder, say "output/" which you can then snapshot, blow away, or do what you like with. Basically though, I keep it separate from data and code inputs.
- aboytsov 14y agoThanks! This much I know. But it doesn't answer my question. Let me repeat it: could you please give me a command to re-build a particular target and everything that depends on it?
- lars512 14y agoThat's exactly what "make -B mytarget" does... Are you thinking of a particular problematic scenario?
- aboytsov 14y agoNo, make -B mytarget rebuilds either mytarget only or mytarget and everything mytarget depends on. A more common scenario is when you need to rebuild mytarget and everything that depends on it. Without rebuilding other parts of the workflow that you don't need.
- tedunangst 14y ago