2 ms·
I work on a distributed runtime system for heterogeneous supercomputers [1]. As an example of the sort of bug we regularly deal with, I am at this exact moment
by eslaught 3y ago
I work on a distributed runtime system for heterogeneous supercomputers [1].
As an example of the sort of bug we regularly deal with, I am at this exact moment tracking down a freeze that occurs on 8,192 nodes of a supercomputer [2]. That means I'm using about 64,000 GPUs and about half a million CPU cores. The smallest node count I've seen my issue is 2,048 nodes and at that scale it only happens about 10% of the time.
We've been debating internally whether Antithesis could help us or not. On the one hand, the fuzzing to explore the state space, and deterministic reproduction, are exactly what we want. On the other hand, we believe our state space is much larger than what you see in a typical distributed database. (And not just because of the sheer scale of things, but even on a single node we have state machines with order hundreds to thousands of states in them.) Based on the post here and the "scenario" count explored in CouchDB, I'm not convinced you'd be able to handle us. :-)
I'd be curious what you think. Happy to discuss here, or contact info in profile.
[1]: https://legion.stanford.edu/ https://legion.stanford.edu/
[2]: https://www.olcf.ornl.gov/frontier/ https://www.olcf.ornl.gov/frontier/
- wwilson 3y agoA drawback of our approach is absolutely that it is expensive to test extremely large volumes of data or compute this way. Even before you start running into physical limitations of our current platform, you will probably be complaining about your bills. :-) Our advice on this is that there are actually a lot of things you can do to exercise behaviors that are usually only seen at massive scale. For example, if you run a distributed storage system, you can probably configure it to split and move shards at 1/1,000,000th of the production size. That might let us hit a tricky codepath much more cheaply. We have a lot more about this in our documentation, e.g. here: https://antithesis.com/docs/best_practices/optimizing.html#knobs https://antithesis.com/docs/best_practices/optimizing.html#k... and here: https://antithesis.com/docs/best_practices/find_more_bugs.html#make-rare-events-common https://antithesis.com/docs/best_practices/find_more_bugs.ht... The other thing is just that the reason many bugs only happen at scale is that they're some kind of subtle distributed race, and you need a lot of nodes for one of the runners in the race to be slow enough that the other sometimes wins. But we can very easily and efficiently provoke these sorts of races by pausing individual threads or freezing nodes, etc. We actually pretty regularly hit issues with tiny deployments that our customers only see in their largest clusters (but no promises, this obviously depends on the details of the software).