3 ms·
Back in 2017 when spectrum and meltdown hit I was boarding a flight to mexico with my wife for a short get away. I noticed a pagerduty as the flight started to
by taf2 3y ago
Back in 2017 when spectrum and meltdown hit I was boarding a flight to mexico with my wife for a short get away. I noticed a pagerduty as the flight started to take off about a backup in our queue. Our company handles millions of calls daily and at that time we indexed each call record as a single document into elasticsearch (ES). When I landed and finally got to our hotel and plugged into the wifi the problem was significantly worse. As a result of the reduced compute capacity from the patches aws rolled out the night before we observed over all servers including our application servers, database and elasticsearch roughly a 30% reduction in computing power. I remember sitting on the beach trying to figure out a good solution. batching the index operations thanks to _bulk endpoint was immediately what I realized we had to do just wasn't clear on how to get there. After 2 days of hacking on the problem and explaining to customers it would be okay no data was lost, I had a solution that involves redis sets to ensure we only index unique call records and a self queueing bulk indexing job that also limits the total number of jobs queued as a function of our overall capacity. Most frameworks assume a single record indexed based on update/create/delete operations is the way to go but that doesn't scale... It's been a few years now since then and elasticsearch has proven it's worth over and over again as a denormalized index allowing for faster search and aggregated reporting that our normalized database could not provide...