8 ms·
I love these articles, thanks! I can't wait for rust to get more involved in this space. How do you plan on doing error handling in timely? Do you expect it to
by erickt 11y ago
I love these articles, thanks! I can't wait for rust to get more involved in this space. How do you plan on doing error handling in timely? Do you expect it to be more along the lines of MPI, where you restart from the last snapshot, or MapReduce/Spark, where lost work is transparently recomputed?
- frankmcsherry 11y agoIt's a good question. I'm punting for the moment because it isn't clear that there is one true solution that makes everyone happy. So, let's say MPI then. :) I think the more optimistic version is probably that the signals Timely Dataflow (the model) gives should make tasteful fault-tolerance logic easier to write. Michael Isard did some work here just before leaving MSR. Realistically, errors haven't been a big problem yet (thanks Rust!) and it's a research prototype for understanding efficient data-parallel compute, so not too bothered yet. The MR/Spark approach comes with lots of hidden costs in the programming model; I'm not clear on when these costs make sense. Using lots more resources so that you survive being killed because you were using too many resources seems like a not-best solution. Some of the layers on top have better fault-tolerance properties. For example, differential dataflow operators are functional and have append-only logs as their internal state; I could see that stack ending up with more transparent fault-tolerance properties, though not yet.