4 ms·
Disclaimer: So do I. Occupational hazard on HN. A year is crazy short if your job is to architect an ideal solution, but insanely long if your job is to simply
by Rule35 6y ago
Disclaimer: So do I. Occupational hazard on HN.
A year is crazy short if your job is to architect an ideal solution, but insanely long if your job is to simply migrate your employer onto a roughly equivalent solution. If a cloud customer isn't continuously testing a migration they're naive.
General best practice is to have a current best setup using all the first-party services on cloud-provider #1 using all their hosted services, and a fallback on cp#2 that self-hosts the services from cp#1 and uses their built-in services where better. It's what you should already be doing for benchmarking and testing but for a fallback if needed.
And once you've got the fallback planned you're free to move the active service around, almost at a whim. Even for a fortune 50 with petabytes of data - if your data isn't already hosted everywhere you're just begging to be wiped out with a simple account problem.
- capableweb 6y agoImagine working for a cloud company and then also saying "If a cloud customer isn't continuously testing a migration they're naive.", not knowing that every user of your software doesn't have time to continuously fiddle with shit that supposedly work correctly at the time.
- Rule35 6y ago> shit that supposedly work correctly at the time. I do love me a well-formed customer query. Yeah, it's supposed to work. I get paid more it if does. But in exceptional circumstances, it won't. And in more exceptional circumstances your alerts will fail too. As an admin what do you do with the 90% of your time that isn't actively fixing something if not planning for how to fix the next thing? We did this years before cloud meant anything other than prepare for floods.
- capableweb 6y ago> As an admin what do you do with the 90% of your time that isn't actively fixing something if not planning for how to fix the next thing? Not every cloud customer has a dedicated admin team, some even don't have any dedicated developer teams. They simply wanted a website done, contracted someone and they uploaded some HTML to a cloud instance. Sure, if you're a large company, it makes sense to have redundancy. If you're running a small company with some esoteric webshop to serve AFK customers, it makes less sense.
- Rule35 6y agoIf it didn't take a team to design what you use then it probably won't take a team to design a fallback. But the difficulty doesn't excuse not doing it, if you don't have a fallback for whatever you pick, pick something simpler. Ideally we keep stuff running, but I don't want to be the kind of person who just tells you that we've got it. You need to know your airbag might fail so that you take your seatbelt seriously.
- dragonwriter 6y agoWhile they haven't actually explicitly said that, and certainly my employer hasn't heard that message, our AWS rep has repeatedly and loudly said that all of our planning should incorporate failure starting from the smallest (like an individual instance) up through a service in an AZ to a service in a region or a complete AZ outage, to a complete regional outage or global service outage to a global provider outage. So, the message is implicitly there.
- chris_wot 6y agoThat's insane. It's ridiculous that you have to constantly test a potential migration!
- andrey_utkin 6y agoIt's not insane if it's rewarded in the long run of the competition. Like, "it's insane to grow legs just to live, if you can just crawl like a worm".
- chris_wot 6y agoThey aren't naive though. Migration testing has costs. A business needs to weigh this against many other costs they have.
- Rule35 6y agoOf course. But you can't just say you can't afford it. Fire doesn't care that you couldn't afford a smoke detector. If you can't afford a recovery plan then what you can't actually afford is the service that needs the recovery plan that you can't afford to develop and test. All costs, including of switching away if it fails, have to be considered as the sticker price.
- znpy 6y ago> If a cloud customer isn't continuously testing a migration they're naive this is a dumb point of view, that only cloud providers (and their employers) can advocate. if i'm a company, being "on the cloud" doesn't automatically makes me money. my business makes me money. if every two years i have to waste a year (or even six months) re-architecting then "the cloud" is costing me people money on top of the infrastructure money.
- dragonwriter 6y agoSo, you architect for multicloud once, and then in the event of a cloud provider failure (a high impact but fairly low risk event if you haven't mitigated the impact with a multicloud architecture) you just shift resources to the surviving providers and maybe throw something on the backlog to incorporate a new provider in your multicloud setup, but even if that takes some rearchitecting, the median interval isn't going to be every two years, or, most likely, even every 10. Most enterprises won't do this, either, but it's not because of recurring rearchitecting costs.
- Rule35 6y ago> if ... then "the cloud" is costing me people money on top of the infrastructure money. Yeah. Do you think I disagree? It's not for everyone. > this is a dumb point of view, that only cloud providers (and their employers) can advocate. No, it's the truth. Stuff fails. And if you think that this is the company's messaging you're totally wrong.
- lazylizard 6y agoyes its your customers' fault.
- Rule35 6y agoSo, pray tell, how would you assess and assign fault here? What's the actual bad action and who caused it? You see, there's strength in the truth. It doesn't matter if drunk drivers shouldn't hit you, that's why you check both ways before crossing the street, and you'd be naive not to. (In this analogy, drunk drivers are outages...)
- dragonwriter 6y ago> If a cloud customer isn't continuously testing a migration they're naive. Most cloud customers are naive, as the stream of major enterprise breaches caused by S3 buckets without security settings that have been default for many years demonstrates.
- ThrowawayR2 6y agoThat's absurd; if a customer has to engineer for and be prepared to migrate to another cloud provider at any moment, that erases any possible cost advantage for using a cloud provider in the first place. They might as well just self-manage bare metal.
- Rule35 6y agoIt's the truth. The truth cannot be absurd. As for running your own servers, that too can fail meaning you still need a migration strategy (even just to new hardware) and you need to be testing it constantly. And no, the benefit of a cloud provider isn't that normal stuff is easy or cheap but that otherwise impossible stuff can be attempted.
- blaser-waffle 6y ago> As for running your own servers, that too can fail meaning you still need a migration strategy (even just to new hardware) and you need to be testing it constantly. That's what DR is for. We have a main bare metal site and a secondary site. Throw a couple of spare servers / switches / PDUs / Hard Drives / whatever in that space too. Cloud options need to be better and more effective than that. > And no, the benefit of a cloud provider isn't that normal stuff is easy or cheap but that otherwise impossible stuff can be attempted. A virtual server is a virtual server. A container is a container. The only thing the cloud offers me is the ability to change my CapEx spends into OpEx spends. Otherwise I have to hope that the vendor won't do me dirty, and will leave me in a stable, workable place 3+ years from now. The bare metal colo operations will. Long track record of stability at everywhere I've been. Barring act of god or otherwise unusual circumstances I know my tier 4 colo will be there next year, and the year after. Will GCP be around?
- Rule35 6y agoWhat's DR other than migrating onto what's hoped to be (but never is) an identical setup? It's like having your Amazon fallback be ... Amazon in another region. That protects you against localized outages but not design failures or systematic outages or incompatibilities in new versions of your stack. And it takes time to keep your DR plan up to date, patch the VMs, etc, and test it. Almost like this migration plan I'm talking about. > The only thing the cloud offers me is the ability to change my CapEx spends into OpEx spends. Ehh, not really. You can setup load-balancer pools larger than your entire colo, or use a globe-spanning backbone to create datasets that auto-replicate worldwide. And which are usually much easier than setting these services up yourself, let alone building the multiple zones. If the cloud is just a big colo to you then you probably shouldn't use the cloud. It's frightfully expensive.