3 ms·
One time, a few years ago a particularly nasty query was executed over and over again and it took a few hours to find it and then block it. And during that tim
by karlney 4y ago
One time, a few years ago a particularly nasty query was executed over and over again and it took a few hours to find it and then block it.
And during that time so many nodes had became slow and unresponsive that another (for us) previously unseen memory leak started to occur.
Nodes kept building up queues of unanswered ping requests on them. And the requests contained our 100Mb large cluster state, so the heaps filled up and evenmore nodes became unresponsive.
And from then on the whole thing turned into a death spiral of doom.
After trying, and failing to get it under control for 48 hours we gave up and rebuilt the whole cluster from scratch, using the snapshots we store on S3.
The recovery took another 90 hours or so. That was not a fun week.