7 ms·
This seems like a non-issue to me. If you're using an IaaS provider you should be treating the network as volatile from the get-go. This is the reason AWS has t
by gdgtfiend 11y ago
This seems like a non-issue to me. If you're using an IaaS provider you should be treating the network as volatile from the get-go. This is the reason AWS has things like auto-scaling groups. You should be designing for failure in "the cloud"
- Fizzadar 11y agoI have to disagree - DO's "cloud servers" are equivalent to virtual private servers, which you would never expect to loose in this manner from other providers. The lack of explanation is what worries me most - it leads me to think this might have been a case of "we forgot to replace a bad drive, then the second in the pair failed".
- bitshepherd 11y agoWelcome to "The Cloud". VPS can be effectively rebranded as "cloud", except when a server has problems you shoot first and ask later.
- RubyPinch 11y agoI'd expect to lose them. Fire happens, floods happen, electrical faults happen, mistakes happen. I can't really expect any human or humans to be perfect.
- e12e 11y agoI remember a ~36 hour down time with a serious uk provider (in my case just a shell account, but they also housed managed and unmanaged servers). They had redundant fiber links to the NOC, from competing providers. But turns out they both ran through the same box under a high way. And then someone blew that box up in order to knock out alarms while they committed a robbery...
- speeder 11y agoMy family business website got killed once by a flood. And many years ago it suffered a truck crash ( the truck crashed on the lamp post right outside the server company building and destroyed the telecom-related wires)
- bluedino 11y agoRemember the blog posts during the early days of AWS? People didn't realize their new elastic cloud servers could go into thin air
- Piskvorrr 11y agoDissolve into thin air? That's what clouds do, right? (One woulda thunk the word "cloud" in itself would be warning aplenty ;))
- __michaelg 11y agoAgreed. This shouldn't be an issue for anyone using cloud services -- or any computer really. Except for some super duper redundant mainframes maybe, you should always be able to deal with losing a component not matter what it is. If you really care, always have multiple instances, in multiple (availability) zones, etc.
- theptip 11y agoExactly, to use a now-trite metaphor: your cloud instances are cattle, not pets.
- cortesoft 11y agoIt is no different if you are running your own hardware - your server hardware can fail no matter who owns the machine.
- lotyrin 11y agoYep, somehow cloud customers have decided that by being in the cloud everything is redundant and durable and has several other magical properties. At one of my positions we had stupidly put too many eggs in the basket of a single physical machine. Its disk controller failed in a way that it trashed the data volumes. I was unable to convince anyone that "move to Amazon" was not a one-step solution to "how do we make sure this never happens again".
- moonchrome 11y ago>somehow cloud customers have decided that by being in the cloud everything is redundant and durable and has several other magical properties To be fair the point of the cloud is that they deal with redundancy, HA, distributing to multiple datacenters, etc. through their services - but you need to use and understand implications of said services to leverage that.
- lotyrin 11y agoWhen cloud == SaaS, sometimes. When cloud == IaaS, nope. But that doesn't stop the misconception.
- techjuice 11y agoThis is so true, I wish there was a way to make this more noticeable (flashing red lights on the order forms, regular email warnings about the single point of failure,etc.). This simple fact is taught in entry level system administration courses (RAID, Disaster Recovery Plans, Replication, High Availability, Fault Tolerance, etc.) but in reality some seem to not plan for it when moving to cloud services. I always suggest people set their apps up to be able to run out of multiple availability zones in case one goes down if it is very important to them. For those that are not able to I will suggest they at least setup replication to another server within the same data center and an offsite location in case their server has hardware failure, an account gets locked or other possible situation to help take of the just in case scenarios that occur to everyone.
- watty 11y agoIsn't "non-issue" a big of an exaggeration? If the dry cleaners lost my clothes, the bank lost my money, a valet lost my car, or gmail lost my inbox I'd be angry.
- chc 11y agoYes, but those aren't expected outcomes when using those services. That's the difference gdgtfiend was pointing out — you should expect cloud servers to sometimes go away.
- bbcbasic 11y agoThey are expected outcomes. Banks fail (think 2008). Dry cleaners do lose clothes. Cars do get stolen. Plenty of stories in the news of all 3 of these. No so sure about gmail losing all your email, but I was able to accidentally access someone else gdocs once. I just logged in as me and saw a complete stranger's docs.
- geofft 11y agoIf the bank lost the particular dollar bill that you deposited last week, but offered you a new dollar bill, would you get angry? If you rent a car from the airport every time you visit a city, and it's always been the same car, but you show up one night and they they tell you the car was in an accident and they'll get you another car, would you be angry?
- AznHisoka 11y agoA more relevant question is: Would you be angry if you lost $500,000 because your bank only insured $150k?
- nkozyra 11y agoThe expectation with these services you mentioned, however, is reliability. The implicit (and explicit) volatility of cloud hosting should change that expectation with such services. "Non-issue" is an exaggeration because it was potentially a catastrophe. But this is sort of the way these services "work." Servers/droplets/instances are ephemeral and replaceable, and their underlying data is not guaranteed in a failure.
- endymi0n 11y agoWhile my first intuition was to agree with you, there's certainly an upcoming generation of developers who have never operated their own root servers and the abstraction level in the cloud nowadays is so high it makes you easily forget that "Droplets" are just VMs are just servers running software. On the other hand, hardware reliability has increased in recent years due to RAID, fully redundant networking & power adaptors, you name it. These facts combined certainly make it easy for the younger generation to forget that redundancy and backup still don't replace each other (never did).
- beat 11y agoWell, there's one way to teach these kids... (Could be worse. Their first lessons could be like mine, in C! If anything will make you paranoid, writing C will.)
- XorNot 11y agoThis seems like the wrong lesson is being taught though somewhere - the delete button is right there. It should be much more obvious that is very easy to make a server go away.
- nailer 11y ago> You should be designing for failure in "the cloud" DO should be designing for failure: VM storage should be on a SAN. A single physical server failing should not cause data loss. This is basic stuff.
- vardump 11y agoThat's a very unreasonable expectation at cloud price levels. Does any cloud provider have SAN backed VMs? If any do, at what price?
- derefr 11y agoThat's what AWS's "Elastic Block Storage" is. You can turn it off and just use instance storage (and I personally prefer to, for truly ephemeral nodes), but it increases spawn time since your disk image actually has to get copied over to the VM host machine in that case, rather than just "attached" over EBS.
- vardump 11y agoThen why EBS failure rate is several orders of magnitude higher than in SAN deployments? A SAN provider would be quickly out of business with 0.1-0.5% annual failure rate. SAN reliability ratings start at 99.999%.
- derefr 11y agoJust because it's a SAN doesn't mean a given abstract block device from it is backed by RAID. It's literally just a multiplexed and QoSed network-attached storage cluster. I actually prefer the lower-level abstraction: if you want a lower failure rate (or higher speed), you can RAID together attached EBS volumes yourself on the client side and work with the resultant logical volume.
- prodigal_erik 11y agoOn AWS, an EBS volume is only usable from one availability zone. You still need to use application-level replication to get geographic redundancy for important data, and when you have that, EBS just lets you be lazy rather than eager about copying a snapshot to local instances.
- nickbauman 11y agoCAP theorem applies everywhere, but most software practices still do not assume that you must code for your app to be able to handle performance variation and partial failure state both of which are common scenarios of cloud infra.