3 ms·
In your post: "There's no more time-honored way to get things working again, from toasters to global-scale distributed systems, than turning them off and on aga
by leghifla 5y ago
In your post:
"There's no more time-honored way to get things working again, from toasters to global-scale distributed systems, than turning them off and on again"
This is generally true, but as all rules, there are exceptions, and I encountered one a few month ago:
To be short, the system (an embedded soft real-time control) ran fine for a long time, and the user added more and more processes. After some glitch and "to be sure the restart fresh", he initiated a reboot... And then nothing worked anymore!
The problem: each process consumed a lot of RAM for a short period at their start. When the user added processes manually, everything ran smoothly. But as soon as a few processes needed to start roughly in sync, it took too much RAM, the OOM killer killed the entire app, and back to square one.
In a way, this is also an example of metastability: the application is restarting in a loop and cannot exit that loop on its own.