9 ms·
My understanding of the frustration is that while the model is deliberately stochastic, there was still unexplained non-determinism in the model that had to be
by jeffbarg 6y ago
My understanding of the frustration is that while the model is deliberately stochastic, there was still unexplained non-determinism in the model that had to be fixed (and still persists). If the randomness isn't fully explained by pre-determined seeds, I'd call that an error.
- tripletao 6y agoHere's the GitHub issue: https://github.com/mrc-ide/covid-sim/issues/116 https://github.com/mrc-ide/covid-sim/issues/116 The discussion isn't terribly clear to me. Is this non-determinism in which random number gets used where, like if you seed an RNG once and then use it from multiple threads that get scheduled in varying order? That wouldn't be good, but the model would still be performing as intended, just with no way to make it deterministic for testing. Or is it some other kind of non-determinism? I'm not sure anyone knows, which indeed isn't totally comforting. I'm not sure the "bug uncertainty" is any bigger than all the existing known uncertainty about the assumptions that go in to his (or any such) model, though.
- Yoric 6y agoDoesn't stochastic imply non-deterministic?
- tripletao 6y agoYes, but when writing deliberately non-deterministic code, it's common practice to make all the pseudo-randomness arise from a single seed value (e.g., a random number generator seed). This is helpful when debugging, since it's otherwise hard to judge whether a change in the output arose from the change you just made to the source code, or from the intended non-determinism. Per my other comment, it looks like Imperial tried to do that here, but it's not quite working. They seem to think the model is fine and just the "deterministic from seed" mode is broken, but hard to be sure until they fix it.
- eesmith 6y agoI don't see anything in that issue which indicates a failure of being deterministic from seed for single-threaded code. Based on the GitHub thread you linked to earlier, there are two issues: 1) multithreading causes problems. Quoting Carmack at https://twitter.com/ID_AA_Carmack/status/1258192134752145412 https://twitter.com/ID_AA_Carmack/status/1258192134752145412 , "issues with random seeds and parallel floating point reductions are extremely widespread, and a great deal of science gets done with non-deterministic simulations." Think of it this way. IEEE 754 arithmetic is not commutative. Not even addition. (A+B)+C might not be equal to A+(B+C). Now what happens if you split A+B+C+D+E+F+G+H+I across three threads. The first compute A+B+C, the second D+E+F, the third G+H+I. The threads finish in the order 1, 3, 2, causing the reduce step to be (A+B+C)+(G+H+I)+(D+E+F). That can very well be different from the single-threaded, left-right addition of those numbers. Hence the reply "We are aware of some small non-determinisms when using multiple threads to set up the network of people and places. (Look for the omp critical pragmas in the code). This has historically been considered acceptable because of the general stochastic nature of the model. We do expect the code to be deterministic when using a single-thread." "omp" here is "OpenMP", a low-level set of compiler directives for multi-threaded programming. 2) The followup reports an reproducibility issue with single-threaded code, which wouldn't be subject to the above problem. Thing is, the reproducible uses the same RNG seed, first when generating a save file, and the second time when loading a save file; both continue from that point. The reporter showed a small difference in the output between the two. This means the RNG in the first case may have been used many times during the creation of the initial data, which didn't happen in the second time. The two generators have different internal states at this point, so should not be expected to produce identical results. Which is why the reply "if the RNG is used in the generation of the network, but not when a cached network is loaded, this might explain the discrepancy. The random sequence in the former case would be advanced relative to the latter."
- tripletao 6y agoFirst, I think you mean associative. That's true, but I don't think that's likely to be the problem here? If the limits of floating point resolution cause such a big change (~80k deaths) in the simulation output, then I'd have bigger questions than the non-determinism. Yes, you can get chaotic behavior solving certain PDEs and stuff; but what in the mathematics of what's basically an extended SIR model could generate a sensitivity like that? I agree that the non-determinism may be pseudo-random numbers used in a differing order, whether from differing seeds or from differing scheduling of multiple threads. Ideal practice would be to correct for these--a thread-safe RNG, an unsafe RNG per thread, seed saved in output files for restart, etc. That they didn't do that doesn't mean their model is wrong, just that it's harder to show that it's right. Note that even in your quote above, they hedged with "might".