3 ms·
Thundering herd has referred to demand spikes in services architectures for at least 8 years[0], probably much longer. 0. https://qconsf.com/sf2011/dl/qcon-san
by mentat 7y ago
Thundering herd has referred to demand spikes in services architectures for at least 8 years[0], probably much longer.
0. https://qconsf.com/sf2011/dl/qcon-sanfran-2011/slides/SiddharthAnand_KeepingMoviesRunningAmidThunderstorms.pdf https://qconsf.com/sf2011/dl/qcon-sanfran-2011/slides/Siddha...
- dragontamer 7y agoHmm, the Netflix presentation there seems to make sense superficially though. The key attribute of the "Thundering Herd" problem is the LOOP. The Thundering Herd causes another Thundering Herd... which later causes another Thundering Herd. In the Netflix presentation, the "Thundering Herd" causes all of the requests to time out, which causes two new servers to be added ("automatic scale up"), then everyone tries again. When everyone tries again, there's more people waiting, so everyone times-out AGAIN, which causes everything to shutdown, add two more servers to scale up, and start over. Etc. etc. Its a cascading problem that gets worse and worse each loop. You solve the Thundering Herd not by adding more resources (that actually makes the problem worse!!), but by cutting off the feedback loop somehow. The problem discussed in the blog post has no feedback loop. Its simply a problem that happens once on startup.
- mentat 7y agoVery good point, thanks for clarifying.
- lmeyerov 7y agoSystem start is indeed a synchronization point, and for limited resources like here, painful and now clpser to the vernacular Thundering herds can cause escalating and successive failures. That is very much an issue with service start/restart. A bad restart will cause a timeout, another restart, and eventually, restarts on further layers. Imagine all this running above k8s. So yes, this pattern is indeed about one of the failure modes that happen with thundering herds. Though if your cache needs another cache, that feels like a bad cache. The promise pattern can be done transparently by the cache, coalescing GETs, instead of requiring a user protocol. We do app level caching to stay process-local because latency is fun in GPU land and we are a visual analytics tool... But that is not for the problem shown here.