7 ms·
North American Object Storage Service Impact
- madsushi 9y agoOnly losing user-uploaded data seems pretty mild, in this case. Meraki sells cloud-managed network hardware, so you don't interact with it that often. Network configurations, logs, traffic data -- any of those would've been much worse to lose. The custom bits like voicemail greetings and IVR will likely be the hardest to replace.
- jacquesm 9y agoWhenever people say they don't need backups because they are 'cloud based' I always wonder what they'll do when their precious cloud provider messes up. The chances of this happening to Amazon, Google or Microsoft are small but they're not 0, if it can happen to Cisco it likely could happen anywhere.
- top_coder 9y agoOne's cloud is someone else's hardware sitting on a basement.
- 5706906c06c 9y agoRemember the AWS S3 outage a few months back? There is a lesson to learn from that.
- NoPiece 9y agoThankfully, they didn't lose/delete existing data during that outage.
- 5706906c06c 9y agoAgree, there was no data loss, but those who relied on US-Standard solely with no bit replication were dead in the water for the duration of the outage.
- Spivak 9y agoSure, but that shows up in the risk calculations when you're choosing a cloud provider. I imagine for just about everyone it was cheaper to eat the loss on the day of the outage than to spend the time/effort/resources to do it right. Especially when it made national news that it was Amazon's fault so nobody blamed the sites that were down.
- jacquesm 9y ago> Especially when it made national news that it was Amazon's fault so nobody blamed the sites that were down. That's an interesting viewpoint. I really don't agree with it though. When your service is down that is your responsibility, never Amazon's. And when you lose data that is your responsibility too, not your cloud provider's.
- Spivak 9y agoWhat was the lesson? That no service has 100% uptime?
- dx034 9y agoAnd that redundancy of data won't help you if the control server crashes.
- org3432 9y agoS3 has had a data loss too sadly, a console UI bug lead to the wrong files getting deleted if I recall correctly. Yet another reason for frontend web tech to improve.
- plandis 9y agoWhen was that? Do you have the post mortem for it?
- org3432 9y agoThere's no public postmortem, it was in 2015 as I recall, it was handled internally and just with affected customers.
- jacquesm 9y agoThere is some proof of this, comment by Scott Bonds, Mixbook: https://www.quora.com/Has-Amazon-S3-ever-lost-data-permanently https://www.quora.com/Has-Amazon-S3-ever-lost-data-permanent...
- org3432 9y agoIt was a bit surprising when I found out, but I think the really interesting part is that low quality web tech are the weak link in the chain. That was an eye opener.
- 5706906c06c 9y agoFactually incorrect, there was no data loss, availability was affected, however; https://aws.amazon.com/message/41926/ https://aws.amazon.com/message/41926/
- org3432 9y agoNot sure why you're citing some unrelated incident, I think you're jumping to conclusions that I'm talking about the one you cited, but I'm not. This one was not made public.
- api 9y agoCentralization decreases frequency of failure but increases its cost. A really large scale nasty incident with a major cloud provider like Amazon could be a national state of emergency.
- jacquesm 9y agoIt really is only a matter of time. The problem with low frequency events is that you never have any idea how realistic your modeling is and the only way you'll find out is when that once in a 1000 year event happens tomorrow morning.
- inetknght 9y ago> once in a 1000 year event happens tomorrow morning It doesn't quite sound like a once-in-a-thousand-year event any more. Rather, it sounds like once-in-a-thousand-year event for a single device, but divided by ten thousand devices means it happens ten times every year. There's whole branches of statistics for failure rates.
- derefr 9y agoI don't need backups of my (S3) backup, because it's more likely my personal backup process has a flaw such that it will backfire and destroy my data, than that it will one day save it from the inadequacy of Dynamo's ~RAID17. Consider: each time you introduce a new device that has local, physical access to the place your data lives, that's one more thing that could Halt and Catch Fire at just the wrong time, or be replaced with a USB Killer or a DMA cryptolocker device by social engineering. If it involves data center operators you don't know, that's more people you have to trust not to break whatever they touch or have been paid off to steal your corporate secrets. Etc. Sure, the probabilities are small—but so is the probability of the great data fortresses crumbling to ash and you being the Last Best Hope for your data. Hypothetical ameliorations of sub-lightning-strike probabilities often have failure modes more likely than their use.
- jacquesm 9y agoIn that case I hope you have your S3 under a different account than your main stuff. There are more reasons why stuff goes missing than just hardware failure. Note that a backup need not make things worse, but should only make things better.
- deleted 9y ago[deleted]
- icebraining 9y agoConsider: each time you introduce a new device that has local, physical access to the place your data lives Right. So don't do that. Put it somewhere else, and configure the original device to push to it rather than give the new device access to the original. You can use a service that implements the S3 API, then you don't even need to install new stuff on the original, just configure an extra endpoint. Also, encrypt before pushing (that counts for S3 too).
- deleted 9y ago[deleted]
- samfisher83 9y agoDo you think you will do a better job of backing up stuff compared to google, msft, etc.? They have dedicated engineers and spend lots of money on this stuff. Think of it from a statistical perspective what is probability of you setting up this back up system vs them?
- wbl 9y agoAddition backups only add, not subtract, from reliability.
- ams6110 9y agoIt's a good point, but I can't help but think of all the mistakes, disclosures, privacy violations, poor design, gratuitous change, etc. that has happened at the hands of "dedicated engineers." In this case specifically, are the engineers who caused this incident not dedicated and well paid?
- icebraining 9y agoThey may have better engineering, but they also have extra risks. My home server will never ban me because it thinks I've violated its TOS, for example.
- toomuchtodo 9y agoNor will your own storage lock you out because you've annoyed a state actor, while a cloud provider will roll over.
- jacquesm 9y agoReasons why data can go missing: - account compromised, wiped out - operator error - malicious employee All of these have happened to companies that I have worked with, so no, I won't do a better job of backing stuff up comapred to google, msft, etc, BUT I would rather have some get-out-of-jail-free card if any of the above should happen and suddenly where there used to be data there is nothing. You should approach this from a cost-benefits perspective, not from a skills perspective.
- 9y ago
- mbesto 9y ago> Amazon, Google or Microsoft Just to be clear, if you're using AWS, GCP, and Azure to host your own applications, you're at your own peril to managed disaster & recovery. Those companies make doing that much easier than managing your own DC and yes the reliability is going to be better than DIY (but still never zero). I think you mean more towards SaaS applications or anything that "phones home" data to back it up, right? We're going to start seeing more business continuity audits of SaaS players, akin to a BBB rating for the company's ability to maintain service levels. I thought I came across a website that actually has started doing this, but I can't recall which it was.
- tgtweak 9y agoI think it's more about shifting blame than actually providing more reliable services. I regularly see s3 throw an error when reading and writing tens of thousands of files (spark w/parquet). It's not that it's MORE reliable (although it is very reliable), it's just that when it isn't it's somebody else's problem and responsibility to fix it.
- creyes 9y agoI designed/deployed a decent sized Meraki network about 4 years ago - at the time it was one of the larger full-stack Meraki networks that exists. An 11 site school district with the edge, all idfs, ap's, phones and at a a later point some of their cameras. Meraki still thinks of themselves a startup, but they have these "uh-oh's" all the time. Random bad firmware that turns off the 5g channel on the MR42's. A DPI "upgrade" that blocked ALL SSL traffic (which at this point is basically all traffic). Their solution was always to try "beta" firmware... in production... in the middle of state-mandated online testing. I was a huge advocate for them but at some point it's gonna be hard for me to keep recommending them. They're so excited about new features but really fail about 1) fixing bugs and 2) ensuring robustness. The "fail fast fail often" mentality really shouldn't work with critical infrastructure
- tgtweak 9y agoUbiquiti (unifi) was very much like this back then too, and to this day still breaks stuff every update. They're getting better now but running a large site or multiple large sites was a constant game if whack a mole trying to figure out which firmware works best on which equipment. I have heard similar stories to yours about meraki and that's what swayed the decision to just go ubiquiti since it's less expensive.
- chrishacken 9y agoWe removed almost every piece of Ubiquiti equipment from our network because of the quantity and severity of bugs in their products. For example: Disabling a port on Edgerouters will grey out the port in the UI, but it doesn't actually stop traffic from passing through the port. I also have no love for how some features are only available via command line while others are only available in the UI. This also differs depending on what product line you're using. Pick one strategy and stick to it.
- tgtweak 9y agoWhat are you using instead now? Genuinely curious.
- bogomipz 9y ago>"The issue has since been remediated and is no longer occurring." I would think that if you lost your data, unless they have restored your deleted data the "issue" is still very much occurring for you as a customer. Wouldn't remediation be that they have recovered your lost data?
- bartread 9y agoThat would certainly be my understanding of remediation.
- rsync 9y agoDoes anyone familiar with their marketing language know how many "nines" they had on their resiliency number ?
- kyledrake 9y agoAs many extra as were punched into that command that permanently wiped out a bunch of AWS EBS data a while back. It haunts me to think about how many people are using these services as their single source of data. Fat fingers melt through 9s.
- SoMuchToGrok 9y agoNot surprised to see this, I've run into many issues with Meraki devices over the past 2+ years. Their support team is amateur at best; at one point I had 6 Meraki engineers working on a DHCP problem (yeah...DHCP) and their recommendation after several weeks of troubleshooting - do a factory reset. I have dozens of stories...don't even get me started.
- oneplane 9y agoAnd this is why outsourcing 100% of your stuff isn't the greatest idea. Sure, the 'managing servers is not your core business'-story holds up most of the time, but when you have no control of anything anymore, you no longer control your services.
- rhizome 9y agoI would think "business continuity" is a core function of, you know, business.
- oneplane 9y agoAs do I, but that's not what marketing teams advertise for ;-) I suppose doing something like specific offloading makes for a great cloud case, but being able to run the baseline at home will at least guarantee a degree of control that will keep you going.
- bartread 9y agoFTA: "Your network configuration data is not lost or impacted - this issue is limited to user-uploaded data." Errrrrrr... so the issue was limited only to data I would actually care about then? Or did I misread? That is a frankly extraordinary use of weasel words.
- wmf 9y agoYeah, you misread. Configuration data is what makes the network actually work and that was preserved.
- bartread 9y agoFair enough. And I suppose that release was meant for people with an understanding of the product, in which case that makes more sense. Still... interesting choice of words. "Oh, it's only the user-uploaded data that we've lost."
- Spivak 9y agoI imagine that from Cisco's perspective the network configuration data is what's actually important. If your priorities are different then I suppose you should be more worried about this incident.