5 ms·
Elastic: Several customer facing deployments deleted across multiple regions
- hericium 4y agoDirect link: https://status.elastic.co/incidents/74yk1h30l5pm https://status.elastic.co/incidents/74yk1h30l5pm
- pertsix 4y agoIs this why the FAA grounded all flights this morning?
- prng2021 4y agoHaha. The FAA isn't using any software developed in the last 20 years.
- nnf 4y agoAutomated deletion is something that has always made me very cautious, to the point that I usually opt to have my code send a notification to myself or the team saying “it’s time to delete x and y” instead of letting the system perform the deletion on its own. In some lower-risk cases, I’ll eventually automate the deletion, but only after a long period of time with no false positives. This strategy has served me well over the years, helping identify edge cases that would have been problematic.
- yabones 4y agoI like the method of adding a "deleted: true" flag to records that are ready to be deleted which makes them hidden in the UI, then log or email something like "INFO: 592 records to be deleted". Then, after another couple days or weeks, a really simple filter removes the "deleted" records.
- plasma 4y agoTo add to this, you can put destructive operations like this into phases, for example, before delete, a power off of the service has to happen, and the delete logic won’t run if the power off time hasn’t elapsed 7 days, etc. These safety layers help present destructive operations more visibly before they are completed.
- cbarrick 4y agoThere are often legal requirements for deletions to occur within a given amount of time. Depending on the volume of such requests, automation is often the only way.
- Someone 4y agoIf you automate everything but pulling the trigger legal time limits rarely will be a problem. For destructive actions, putting a few days between “take server offline” and “throw disk into the shredder” often is possible, too.
- abujazar 4y agoAnd this is why I don't rely on cloud services for mission critical applications.
- hdjjhhvvhga 4y agoIt doesn't matter if it happens or not. What matters if they manage to recover and how long it takes. Based on that information you can make meaningful decisions regarding the risk. I can't imagine I'd put all eggs into one basket these days.
- leros 4y agoI think about it exactly the opposite. I'd rather have a team of engineers at Elastic waking up at 3am and fixing things than me having to wake up and fix things by myself. My most critical infrastructure for my one-man SAAS is all third party infrastructure run by large companies. My non-critical infrastructure is self managed for cost savings.
- abujazar 4y agoSo I guess my comment is getting downvoted by people who don't understand «mission critical». >how can you run mission critical if most cloud services have SLAs two or three nines weaker than needed That's exactly it. Cloud providers usually provide SLAs in the range of 95-99%. Amazon doesn't provide a full refund until monthly uptime goes below 95 %. Elastic apparently doesn't provide an uptime guarantee at all, they only provide an SLA for support ticket response times. And only on gold and platinum subscriptions. This incident lasted for more than 24 hours («most» instances restored after 22 hours). It doesn't matter if it's someone else that has to wake up at 3 am to fix the issue, when they're unable to fix it within reasonable time. Mission critical apps simply can't be down for 24 hours. And it's fully possible to design HA elasticsearch deployments. Elastic Cloud just isn't one of them.
- aenis 4y agoSadly, doing things in house is not inherently safer. Same type of people designing same type of processes. But yeah, how can you run mission critical if most cloud services have SLAs two or three nines weaker than needed?