6 ms·
I think this adds some momentum to the pendulum swinging back the other way. Maybe cloud teams can patch your services better than your in-house team can (See 2
by SCHiM 5y ago
I think this adds some momentum to the pendulum swinging back the other way. Maybe cloud teams can patch your services better than your in-house team can (See 2 critical issues in Azure the last 3 months, _caused_ by MS itself).
Maybe the cloud has a higher uptime than your on-premise infrastructure (see the AWS, Azure outages). Make sure to compare the actual outage time v.s. the stats doctored by various political pressures and weaselly worded SLAs (how do you mean you had an outage? Only 49% of your requests were failing!).
- tinalumfoil 5y agoEven if you can manage more uptime on your own than through the cloud (which I doubt), being on the cloud means downtime is correlated with downtime of other services. That's usually a good thing. Your customers will be more understanding if your outage is part if a wider outage that makes national news. Any services you integrate with are likely down too. If two services with 99% uncorrelated uptime together drops to about 98%. It doesn't drop if the downtime is perfectly correlated. Even if you don't directly integrate, your customer's workflow might. They see many services down and say, "cool, time to get caught up on laundry". If only you're down it's more aggrevating.
- 5e92cb50239222b 5y agoWha? We host everything on our hardware (which is nothing special) and haven't had any downtime in this year (yet). And we're just another run of the mill dev shop, very far from "superstars" who work on these (supposedly extremely stable) platforms.
- oxfordmale 5y agoTwo questions I have regarding your in-house hardware: 1. How easy can I access your physical servers ? 2. What happens if there is a catastrophic failure, for example local power outage or a major flooding 3. How secure is your server? Are you regularly patching your operation systems 4. If I want to run a project that requires double the capacity of your current hardware for a specific project, how long is it going to take to get it spun up?
- christophilus 5y agoI'm running about 10 servers myself in production. They just do transcoding, so aren't mission-critical, but... - Regularly patching is automated and took about 30 seconds to configure, using an automated script Regarding running your own physical servers, that is a different ballgame, but for all of my projects, if I need to: - I can pretty easily spin up VPSs / bare metal servers anywhere (netcup, linode, hetzner, etc) and provision there while I wait for new hardware to come in - If you want to double the capacity of your current hardware, you'll have to order it and wait, but it's cheap (vs the major cloud providers) to way over provision if you're running your own physical hardware, so you can pretty easily have 2-4x extra capacity and still come out with extra money in your pocket. I host in the cloud, but I think people vastly over estimate how much it saves 90% of cloud customers.
- grey-area 5y agoThere are many approaches which don't depend on AWS, and not all of them mean hosting your own physical servers, and they certainly don't mean you don't have an off-site backup policy. There are well-understood answers to all your questions, they are not too difficult, they just cost money - some businesses choose not to spend that money, some weigh cost-benefit and go for AWS, some decide to go for in-house servers, some go for hosted Virtual Servers, some go for serverless. Why does everything have to be built one way?
- gregmac 5y ago> they just cost money - some businesses choose not to spend that money, some weight cost-benefit and go for AWS, some weigh cost benefit and decide to go for in house, some go for hosted Virtual Servers, some go for Serverless. I'd also say some choose not to spend the money, but fail to consider the cost of that choice. For example: Doing old-school manual deployments that require herculean efforts to update at off-hours on the weekends burns people out, and makes it hard to attract new talent. In other words, you've made the decision to spend more money on finding and retaining people. But it's definitely way cheaper to pay for colocating a single Dell server you bought 3 years ago than what you'd spend in the same time on AWS. And if your hardware never dies, paying for the redundancy might seem silly. A lot like paying for fire insurance despite the fact your house has never even burned down.
- throwaway984393 5y agoWhen I worked as an IT person for a cheap hotel back in the day, I set up a single Compaq PC as the only server (Samba NT DC, file sharing, IP masquerading, web/mail/DNS server, etc) in a closet. It was not on any battery backup, there was no modem. It ran for years without ever rebooting. In a city where power outages were normal. Similarly, I've had EC2 instances run for years without ever being rebooted or going down. One of them's still running today after 5 years. But none of those services were being used 24/7; if the internet went down for an entire weekend, or a hard drive was a little bit corrupt but kept running programs in memory, I probably would never have noticed. I've also had EC2 instances literally just fall off the map and sort of disappear, and had them manually replaced without notice by AWS, had virtual drives fail and corrupt, and had calls to services fail. And I've had my desktop's power supply get fried by a power surge. Without a lot of experience, running systems seemed easy. But as time went on I learned that it can be easy, and it also can go down if somebody blows on it the wrong way. What we see as being reliable may just be chance. The only way to guarantee reliability is to expect that things are going to go down, and design and build it accordingly. What AWS makes easy is they give you all the components for reliability, but you have to do the plumbing yourself. I'll bet you the people whose services went down did not properly design for reliability, as they were probably running in one region, in one set of AZs, and relied on distributed system operations that can fail, and didn't properly account for how to deal with those failures. One product I maintain on AWS did not go down, but another did. Also, the more components a system has, the higher probability there is of failure. Big systems are actually more error prone than small ones.
- justbaker 5y ago> Your customers will be more understanding if your outage is part if a wider outage that makes national news. Yes they definitely are.
- newaccount2021 5y agoExactly! Furthermore, your customers will likely be AWS customers themselves so will have bigger issues to contend with
- savant_penguin 5y agoOr they could see their entire product line break at the same time, in such a way that they cannot mitigate
- ocdtrekkie 5y agoMy desktop PC has better uptime than the cloud right now. So does my datacenter. In fact, I am not sure I have anything electronic that doesn't have better uptime than AWS or Azure.
- y4mi 5y ago> Even if you can manage more uptime on your own than through the cloud (which I doubt) Getting higher uptime is super easy for smallish inhouse deployments. You just don't install any updates and let the server run, trusting your VPN to shield you from possible security issues. The maintenance burden is the reason why people often prefer the cloud services, not the uptime. Because maintaining the instance with updates, reading all patch notes and steps for migration, keep every health metric monitored and respond quickly on issues without getting stuck googling for possible reasons is quite a bit of work and quickly forces you to employ n+1 people.
- api 5y agoThis is particularly true with modern fully solid state hardware. Uptimes of decades are likely possible if you don't mess with anything. We run some bare metal servers. They just never go down. Solid continuous pings for years as monitored from elsewhere on the Internet. That's because they're just boxes on a rack somewhere running an OS and some steady-state services (ZeroTier roots). Simplicity is more robust than complexity. SaaS is definitely about the pain of managing and upgrading software, but it's also about OPEX vs CAPEX. Many companies will pay more for things to put them in the OPEX column for various entirely synthetic accounting, investor relations, and tax reasons. I do wonder if the pendulum there will swing back though since if you price out cloud vs. physical hardware the market has become extremely distorted. Many companies spend enough on AWS to buy an entire rack of hardware at a different data center every month and pay 2-3 employees to manage it. That hardware would be up to 100X as fast and powerful as what they rent at AWS and bandwidth would be almost free. That's a really distorted market. The amortized costs should not be this different.
- carlmr 5y ago>Many companies spend enough on AWS to buy an entire rack of hardware at a different data center every month and pay 2-3 employees to manage it. That hardware would be up to 100X as fast and powerful as what they rent at AWS and bandwidth would be almost free. That's a really distorted market. The amortized costs should not be this different. I think the issue here is that it's not zero-cost to switch. Your processes will adapt to some implicit assumptions that aren't true outside AWS, Azure, or whatever vendor you locked yourself into. If we somehow managed to have a completely standardized interface here, the market would be more competitive.
- SCHiM 5y ago> That's usually a good thing. For any individual company able to offload the blame, that's great. It's not so great if half the countries' doorbells, robot cleaners, various home streaming service setups, the baby camera, the fridge, the smart TV and your phone stop working... All at the same time. In my opinion, the 'downtime' really should be measured in $NUM_SERVICES_STOPPED X $TIME, instead of just $TIME. And in this case I think any long time Amazon outage is orders of magnitudes worse than your regular old slow IT company outage.
- albertopv 5y agoAn our on premise linux server recently reached an uptime of 1000 days. Yes, days.
- jve 5y agoDo you not patch your servers? Or you don't need reboot for Linux devices?
- heyitsguay 5y agoCan't speak to their particular config, but at least LivePatch for Ubuntu can apply most updates without the need to restart.
- albertopv 5y agoI dont know, we received a happy email from our IT team to celebrate 1000 days of uptime just few days ago.
- salawat 5y agoMost globalist thing I've heard all week. Make things that fail with everyone else. Your customers will just have to get over it because everyone else is down. News flash: This is exactly why "overconsolidation" is a bad thing. The bigger, more complex, and integrated the player, the more devastating the eventual failure is as reality dictates you will build the most complex system possible until you outstrip your ability to mentally simulate, reason about, and debug it. You cannot create a thriving, resilient business ecosystem With everyone flocking to the same players. Customers must come first. Not your books. Working, resilient solutions. Anything else is LARPing, and kicking the catastrophe can down the road to someone else to pay.
- jve 5y agoIf you happen not to be on the major cloud platform where everybody is enjoying donwtimeshare, we provide you with some downtime monkeys that will pull the lever EXACTLY when your platform should go down! Let no more have your integrations with systems in the cloud frustrate your customers when they are unreachable - let the news explain the downtime. JOIN the Downtime Umbrella NOW and receive 5 downtime lever pulls for FREE! /s
- ipaddr 5y agoWhy would I want to be down when everyone else is down? That's a period where I can get customers from everyone else.
- bob1029 5y agoI'd rather have control over my circumstances than the ability to assign blame for them. "my outage means the customer is probably also down so they maybe don't care" is not a viable way to run a business in my mind.
- gerbilly 5y ago> being on the cloud means downtime is correlated with downtime of other services. That's usually a good thing. A long time ago, I used to circulate snarky little emails at work. One of them was responding to this very concept. My managers were throwing out our working and mature UNIX servers (implementing DNS, Mail, and other services) in favour of NT. The new system crashed a lot, we had some security breaches, but at least it was 'industry standard.' Managements' response was that with the UNIX stuff we had no one to pin our outages on. Now we could blame Microsoft, and call their support line. I circulated an email making fun of this justification, which promoted a fictional product called 'Blame Studio' which would help you map out the blame path for any of your products or services. It would help to make sure that none of the blame ever landed on you, but rather was always redirected onto some other company.
- nabla9 5y agoAlso, not all uptime is equal or worth of the price. What would be more valuable to you: 1. 99.5% uptime where unscheduled downtime is max 10 minutes vs. 2. 99.5% uptime where unscheduled downtime comes in 2-5 hour chunks.
- ocdtrekkie 5y agoAnother key point of this: Selecting your downtime. If you own your own stuff, your major changes and maintenance are done at times favorable for your business. AWS or Azure configuration fat-fingers happen when best for their business, not yours.
- gtirloni 5y agoI don't see any momentum in that direction at all. If anything, it's quite the opposite. These outages have almost zero impact on anyone deciding to use a cloud provider.
- csw-001 5y agoAgreed. In fact, I've heard government decision makers actually swayed to go to the cloud due to outages with logic along the lines of, "If Amazon can't keep AWS running, what chance does our 30 person, chronically underfunded, team have?" There is also safety in having the outage be somebody else's fault, and having that outage effect a huge swath of the economy - "Sorry customer, but it's not just us." Reminds me a bit of the "nobody gets fired for picking IBM" mentality.
- tuldia 5y agoThe learned helplessness in the cloud is stupefying, so many outages and downtime that could have been avoided by a competent admin.
- spamizbad 5y agoOne problem facing on-prem orgs today is the sad state of commodity server hardware; poor quality control, buggy firmware (and vendors who won’t help), BMCs with massive attack surfaces (that have already claimed one VPN provider), and commodity network switches that can push lots of packets but aren’t terribly flexible for people who are pushing more challenging payloads around their DC (video, etc). "Hyperscalers" like Amazon, Microsoft, Facebook and Google build their own hardware and are able to avoid many of these problems. Unfortunately, none of this stuff is available off the shelf to mere mortals. There’s a startup trying to fix this problem (Oxide) which I think is launching their racks next year. Will be interesting to see what happens.
- jjav 5y ago> Maybe the cloud has a higher uptime than your on-premise infrastructure (see the AWS, Azure outages). Not even so sure about that. I've had a ton more downtime ("degraded" in AWS speak) with AWS than any self-hosted systems. And that's with more than half my career on self-hosted. If a major disaster strikes, like the whole rack catching fire and melting everything, then it's true that AWS could recover quicker than self-hosted. But most problems are not of that sort.