7 ms·
You are right that makefiles alone would be a bad fit for that situation. My approach would probably be to wrap the parameterizable parts in a script accepting
by xaa 12y ago
You are right that makefiles alone would be a bad fit for that situation. My approach would probably be to wrap the parameterizable parts in a script accepting arguments and run the combinations with GNU parallel. Possibly controlling the overall execution and dependencies in a makefile.
It's less "clean", because you have to keep track of metadata like parameters in the output filename or similar. But the advantage is a huge increase in flexibility.
We tried Celery for awhile as a job manager which looks like it has similar capabilities to Luigi. It was slower, caused us to write a lot of ugly "bash-in-python" when calling non-Python programs, and broke the UNIX philosophy of having independent programs doing one thing, making it harder to quickly test new combinations of components without writing a lot of Python code.
It also depends on your dataset size. We do a lot of machine learning on datasets that won't fit in RAM, which is perfect for the pipe/streaming model.