10 ms·
DigitalOcean block storage is down
- kyledrake 7y agoWhat unholy thing did they do that broke it across 12 different datacenters, good lord.
- bluedino 7y agoProbably the old "one command ran on everything"
- astrodust 7y agotmux is a dangerous tool in the wrong hands.
- rubbingalcohol 7y agoto be fair, it's dangerous even in the best hands. mistakes happen but business processes need to be in place to prevent catastrophes... every time i see something like this, my inclination is to blame the CTO, not the engineer who pulled the trigger.
- toomuchtodo 7y agoA post mortem should always be a place to highlight deficiencies in processes and communicate necessary improvements put into place, not to blame. Blame should only occur if the cadence of outages becomes excessive. Complex systems are tricky, and to err is human. Disclaimer: Ops/infra engineer in a previous life.
- astrodust 7y agoI wonder how many outages these days start with something like "kubectl apply" and then things go horribly awry.
- deleted 7y ago[deleted]
- solotronics 7y agoWe can blame whoever we want but you better believe shit rolls downhill at most places.
- GhettoMaestro 7y agoUntil it is a big enough F-up that an executive's head must roll.
- mdaniel 7y agoThere's a famous corollary to that approach: "Fire him? Why, I just spent 10 million dollars _educating_ him" (regrettably I can't find any evidence it's a true (quote|story), but I enjoy the sentiment)
- alexeldeib 7y agoThis does seem to indicate a notable lack of isolation for the blast radius between DO datacenters. Would be interesting to see the post mortem.
- protomyth 7y agoI get the feeling that whoever writes the post-mortem is going to have a bit of pressure to assure folks that there is isolation going forward.
- swsieber 7y agoWhoever broke it is going to feel significant pressure to actually isolate things too.
- klodolph 7y agoThat would be a bad sign that there’s something wrong with the culture. I would hope for a postmortem that identified flaws that genuinely needed to be fixed.
- viraptor 7y agoThose are not mutually exclusive and actually a good idea. You want to fix this specific issue, but also ensure that whatever process took down one DC doesn't affect other DCs. That's scaling and redundancy 101 - not sure why it would be something wrong.
- klodolph 7y ago> Those are not mutually exclusive and actually a good idea. The goals “assuring folks that there is isolation” and “identifying flaws that need to be fixed” are somewhat contrary to each other. The post-mortem should identify flaws in systems, processes, and thinking. It should not try to assure people that there is isolation when there is evidence to the contrary. > You want to fix this specific issue, but also ensure that whatever process took down one DC doesn't affect other DCs. This was a multi-regional failure. So, this specific issue is also an isolation problem, among other things. You will want to ensure that this problem doesn’t happen again but you shouldn’t assure that it won’t.
- mdellavo 7y agodee ennn esss
- hinkley 7y agoA bug that has no obvious side effects that only became visible once all data centers were upgraded? Happens. Statistics are hard.
- usernametologin 7y agoThankfully they fired the VP last year who is a complete idiot when it comes to infrastructure. Probably related to decisions he made no doubt. Companies prior to DO are still fixing his bad decisions. And now he's in charge of blockchain stuff. wcpgw????
- pmlnr 7y agoBad puppet/ansible/etc commit is the most probably explanation.
- nonbirithm 7y agoIt could be DNS. Azure has had an all-region failure due to a single DNS provider outage. It was possible that same DNS provider's outage was also causing problems for GCE and AWS at the same time. https://news.ycombinator.com/item?id=19812919 https://news.ycombinator.com/item?id=19812919
- dc352 7y agoThat wouldn't be at the top of my list. We have "Volumes" for databases and they were inaccessible for like 6 hours. I don't think any DNS is involved in mounting these. But hey, there's always a lot of crap hidden behind the scenes :)
- markonen 7y agoI would be absolutely amazed if DNS was not involved in mounting a block storage volume.
- golanggeek 7y agoThis is really down for more than 2 hours!!!
- sondh 7y agoLast night I was testing DO managed Kubernetes cluster with persistent volume claim and the volume took 15 minutes to reattach after the pod is rescheduled to another host. I thought it was just some weird hiccup and went to bed. The incident report indicated the problem started 4 hours ago (around 9pm GMT) but I was having problem around 4pm. It's definitely not a 2-hour incident.
- dc352 7y agoour disks in London went down at about 8:45pm UTC (10 mins 100% disk utilization alert triggered at 5 to) and DO recovery message was sent out at about 2am UTC. We switched our service (keychest.net) on at 3:15am
- hartator 7y agoNot sure why the previous incident page got flagged. This is the new one. It's affecting us for real. Making almost our whole service - serpapi.com - down. As we are storing database files on block storage volumes.
- dang 7y agoI took a look at the flags on these stories and am pretty sure they're from users who are tired of "X is down" submissions, which tend to get posted a lot and often to be a little on the trivial side. However, since several HN users are expressing that this issue is genuinely affecting them, I've turned off flags on the OP about this and merged the comments here.
- tyingq 7y agoThe "across all regions" part makes this one different for me, and interesting even though I'm not a customer of their block storage. I'm curious about the sequence of events, or design choices, that would cause that.
- astrodust 7y agoI reported it and they were like "what? oh..." Then the status page changed and as things got worse, the dashboard page got an announcement as well.
- CaliforniaKarl 7y agoPersonally, if DO don’t have anything new in a status post, I’d prefer seeing an update that says something like “We are continuing to work on the issue. Nothing new to report. Next update in X minutes.” That is a lot easier for me to parse than the text that someone seems to be copy/pasting in each update.
- iamsb 7y agoWould be great if statuspage.io has a button when pushed publishes message similar to your suggestion.
- imglorp 7y agoHrm, Atlassian BitBucket is also down. Just a coincidence? Does BB use DO? https://bitbucket.status.atlassian.com/incidents/4t1pkwrdtl8b https://bitbucket.status.atlassian.com/incidents/4t1pkwrdtl8...
- iamaelephant 7y agoBitBucket definitely doesn't use DO.
- privateSFacct 7y agoHigher latency (per status) is not end of world especially if it’s just “may experience” higher latency.
- erikrothoff 7y agoThat wording kinda ticked me off because our volume was completely inaccessible. Rebooting did not mount it at all.
- sb8244 7y agoIt looks like they have just updated it as resolved and monitoring.
- sunasra 7y agoI was always wondering how I can get know proactively if something like this break or some service has an outage. As a result, I have built this tool( http://incidentok.com http://incidentok.com )
- sunsetMurk 7y agogreat idea - going to give it a whirl this week. i'm curious about the slack integration. can you provide some more info on what that looks like? eg. just a message in real-time when it goes down? a daily message of statuses? etc. Any sort of customization w/ it? I currently use a soup of zapier zaps to take care of this problem.
- sunasra 7y agoHey. Thanks IncidentOK will send message to slack using webhook as soon as incident reported by any product. Didn't thought to send status everyday. But I am open for suggestions Message looks like this https://imgur.com/jjbMKj8 https://imgur.com/jjbMKj8
- simplehuman 7y agoAnyone have a review of using DO k8s or DO managed DB in production?
- stephenr 7y agoThis is your weekly reminder that anything you want to be reasonably “HA” should span multiple vendors in multiple DCs.
- Bombthecat 7y agoYeah, the myth that "just use aws" to have 99, 9999999 percent uptime is coming to an end...
- stephenr 7y agoOh I'm sure the myth will persist for many years.
- dc352 7y agothat would be pretty cool but to have that, you need a high-network-latency solution, i.e., pretty much cold back-up. For some time I thought it's pretty last century option but having been experimenting for some time now, it's the option with lowest impact on system performance. More importantly, it's reasonably resilient.
- stephenr 7y agoI've read your comment now about 4 times and all I have come up with is "huh?" Literally thousands if not millions of organisations operate multi-DC infrastructure across the planet. Is it harder than setting up a single box in one DC? Yes. Is it harder than setting up a mini-cluster of boxes in one DC? Yes. Is it rocket science? No.
- seaghost 7y agoTheir block storage is such a failure. I’m back and forth with support to automatically delete files with lifecycles for over 2 months now and it’s still not resolved.
- ngrilly 7y agoSince you're trying to "delete file with lifecycles", I'm quite sure your problem is with their object storage (called Spaces), and not their block storage.
- jacquesm 7y agoThank you Digital Ocean for once again proving that 'The Cloud' is not a backup.
- vinw 7y ago'The Cloud' is _a_ backup. Just don't let it be your only backup!
- pnutjam 7y agoIt's not backed up if you don't have 3 copies.
- lunchables 7y agoThe saying we always use is "If it is not in 3 places, it doesn't exist." And another: "3 copies, at least one offsite"
- louwrentius 7y agoIsn't Digital Ocean running Ceph for their block storage? I would wonder - as others suggested - that they may have stretched the cluster across datacenters ?! Would be interested in the post-mortem.
- ngrilly 7y agoYes, DO uses Ceph: https://blog.digitalocean.com/why-we-chose-ceph-to-build-block-storage/ https://blog.digitalocean.com/why-we-chose-ceph-to-build-blo...
- jbverschoor 7y agoTheir ad was “you’ve been developing like a beast and your app is ready to go live” DO is a nice thing to play around with and maybe launch something, but I wouldn’t run full production on it.
- pastrami_panda 7y agoThis is OT, but I have a droplet on DO and I'm amazed at the amount of malicious traffic it gets. Is it normal for a very private vps to receive thousands of ssh attempts per hour? I have fail2ban installed and the jail is so busy it's quite astounding. Anyone with more web hosting experience that can weigh in?
- pmlnr 7y agoIt is "normal". Even my home fix IP gets it without any service running on it other than ssh.
- zeta0134 7y agoI work for a web hosting company in Texas, and this is ridiculously common. Any public IP with any public service at all will be poked, prodded, and generally made uncomfortable by every bot and crawler you can think of, trying common password combinations and scanning for common vulnerabilities in popular software. This catches so many of our customers by surprise, who tend to mistakenly believe they're being targeted in some kind of attack. Generally they're not, unless they're running something vulnerable and one of the bots noticed. Fail2ban is great to at least stem the tide. It's good at slowing down SSH brute forcing, and can be set up to throttle poorly behaved scrapers so your site isn't getting hammered constantly. If you can deal with the inconvenience, it's even better to put services that don't need to be truly public behind an IP whitelist. That stops the vast majority of malicious traffic, most of which is going after the low hanging fruit anyway. Otherwise, it's kinda just a fact of life. With the good traffic also comes the bad.
- pastrami_panda 7y agoCheers for weighing in. A whitelist is a good solution, since the sheer amount of attempts is making me uncomfortable. It seems to be accelerating over time as well which is even more disturbing.
- davrosthedalek 7y agoI always switch my outward-facing ssh servers to key-only. Is there any advantage for running fail2ban additionally?
- sodosopa 7y agoSo that’s why bot attacks and spam traffic was lower.
- irfanbaigse 7y agoDigitalOCean bad experience
- unilynx 7y agoDigitalOcean just posted a post-mortem on http://status.digitalocean.com/incidents/g76kgjxqrzxs http://status.digitalocean.com/incidents/g76kgjxqrzxs (the same url)