4 ms·
On first read this looks almost completely useless for actual distributed systems. So first off, distributed systems use different software for different tasks
by exfalso 3y ago
On first read this looks almost completely useless for actual distributed systems.
So first off, distributed systems use different software for different tasks. These will not be/were not written in Rust, and even if they were, they may not use Tokio (which is a library I could go into a rant about as well...).
Second, the very premise of the library, namely that it simulates the components and serializes everything into a single deterministic thread, is already wrong. Distributed system issues are difficult because they are non-deterministic, the faults happen rarely, and they exhibit Heisenbug behaviour: the more diagnostic you add to the system, the more likely you are to serialize the event sand therefore hide the bug. This library starts off with serialization, effectively hiding most if not all issues that can happen in a real system. What is the point of this?
So how would an actually useful distributed testing framework look like? It would be almost the exact opposite of turmoil:
1. It would handle an actual distributed system with different components written in different languages
2. It would inject payloads that are likely to tease out races/consistency violations
3. It would inject failures at various layers (network, storage etc) to test how the system recovers
4. It would provide useful diagnostics, and a way to reduce the failure scenarios to the extent possible (again, we're talking about non-deterministic failures, so this is not easy)
- sagarm 3y ago> distributed systems use different software for different tasks. These will not be/were not written in Rust, Obviously, that's not always true. Having a way to do deterministic testing in a single process is helpful. > Distributed system issues are difficult because they are non-deterministic, the faults happen rarely, You do not want non-determinism in your tests just because the real world non-determinism. How else are you going to write repeatable tests of scenarios involving concurrency? You can build fuzzing on top of a deterministic simulator, allowing you to reproduce failures. This is a strategy that FoundationDB has used to great effect. https://apple.github.io/foundationdb/testing.html https://apple.github.io/foundationdb/testing.html
- exfalso 3y ago> Obviously, that's not always true. Having a way to do deterministic testing in a single process is helpful. It is, but then you're not testing a distributed system. > You do not want non-determinism in your tests just because the real world non-determinism. How else are you going to write repeatable tests of scenarios involving concurrency? You can't, but you can get close. E.g. with fuzzing you can generate the scenarios with some ordering constraints, so if you replay the scenario a couple of times you are likely to reproduce the failure. It's a balancing act, too much serialization and your tests are useless, too little serialization and the state space becomes too large.
- exfalso 3y agoReading about it, the FoundationDB testing does look very good. If you control all of the atomic events and build a fuzzer on top, you can indeed end up with reproducible scenarios testing concurrency, up to the atoms' scope. In the async scenarios as with Tokio and Flow(which I just learned about, cheers), the atomicity extends to the scheduler yields, which is already pretty good. It cannot test e.g. atomic memory operations/memfences etc but it can test in a fairly fine-grained way. However, my original point still stands: you can only do this if you control all of the system within your particular scheduler. In real world scenarios you are more likely to encounter: 1. Some database written in C 2. Some storage layer maintained by AWS/Azure/GCP 3. Some messaging layer written in Go 4. Failover modes of the Kubernetes/Hashicorp stack 5. Your node.js webserver named "backend" 6. Some performance-sensitive component written in Rust This sort of testing just does not apply