4 ms·
I worked with long (>> 24 hours, some times up to a week) complex workflows on big (thousands of nodes) clusters. We used custom software layered on top of a jo
by fpierfed 8y ago
I worked with long (>> 24 hours, some times up to a week) complex workflows on big (thousands of nodes) clusters. We used custom software layered on top of a job scheduler like PBS Pro or HTCondor. The nice thing about this setup is that it supports re-running failed jobs, has pretty good monitoring, does an OK job at resource selection and allocation and is language agnostic. The last point is good if your workflows have parts written in different languages. There are a handful of conferences a year on these topics by the way. My favorite is HTCondor Week at the University of Wisconsin in Madison. Talks are online [1]
[1]: http://research.cs.wisc.edu/htcondor/HTCondorWeek2017/ http://research.cs.wisc.edu/htcondor/HTCondorWeek2017/