13 ms·
GCP Incidents
- esafak 3y agoWhen they say moving are off Google Cloud services to bare metal, where do they plan to move?
- b112 3y agoMy response to this, is that there are endless ways, and places to do this. There are far more colos, people that will rent you a rack, and bandwidth, than VPS types. And you can rent servers too, instead of buying your own. Colo is literally 10000x cheaper than many AWS deployments. I've seen million dollar bills drop to tens of thousands per year. And of course, you can always deploy in house, in your own server room.
- bagels 3y agoSure, colo can be cheaper. What kind of infrastructure was this? Or was this just bandwidth bills?
- deleted 3y ago[deleted]
- dilyevsky 3y agoI’ve done some modeling and for our high cpu, high egress, medium storage multi-million $ a year deploy it was 70-90% lower cost than cloud when you factor in amortized cost of boxes, remote hands, transit etc. Pretty substantial but not 10000x ;)
- iot_devs 3y agoWas this comparing AWS on demand pricing or the 3 years plan?
- dilyevsky 3y ago3 year plan. Basically for colos you also get it cheaper if you sign for 3y so that’s more apples to apples. Equipment “annual” cost was calculated over 5-7 year lifetime - I didnt go as far as calculating how much you could recover if you pawned it after 3 years…
- pseg134 3y ago[flagged]
- dang 3y agoWe've banned this account for repeatedly breaking the site guidelines. If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html.
- tehlike 3y agoMy guess is hetzner.
- philipswood 3y agoHetzner is awesome, but last time I checked it's consumer-grade hardware.
- supriyo-biswas 3y agoMany data centers provide colo/hardware renting facilities, such as Equinix, Coresite, Digital Realty etc. (Even AWS got started off those, though they mostly build their own data centers now.)
- londons_explore 3y agoWhen small companies get big, there must sometimes be legacy compute jobs still left on the original infrastructure right? Ie. Jeff's original php script to trigger some automation that never got adopted by any team. Do all the big tech companies still have a box somewhere full of legacy stuff that 'probably isn't important, but not worth turning off just incase it is'?
- ur-whale 3y ago> where do they plan to move? Basement of their office? We reached the same conclusion they did a while back and went back to good-old self-hosted. Reliability has been as good as cloud and TCO is divided by a factor of 10.
- ghusto 3y agoA data centre or (less likely) their own office. This was the way things were done not that long ago ;)
- asylteltine 3y agoAnd they are terrible for ops and security. People also used incandescent lights not that long ago
- dharmab 3y agoAny business can rent space in a colo pretty easily. The constraint is mostly hiring engineers with experience racking and stacking boxes, and willing to drive to the colo when on call.
- 363082a9-58a7 3y agoI've had an experience with GCP that involved a very enterprise-y feature breaking in a way that clearly showed the feature never worked properly up until this point (aside from causing downtime when they tried to quietly fix it). GCP reps proceeded to remind everyone in the call in which they were supposed to explain what happened they were under NDA, because admitting to the above would've been a nightmare for regulated industries.
- HenryBemis 3y agoI always wonder whether an NDA can prevent you speaking/whistleblow to a Regulator, Police, DA, or some (truly) state authority. I would like to assume, 'no, you can always report a crime'.
- supriyo-biswas 3y agoCompany lawyers like flexing their muscles regardless of the actual legality of such agreements or clauses. As an example, in India, non-competes are outright illegal since the Indian Contracts Act directly states that any clauses in a contract restraining a lawful profession or trade will be disregarded, and yet most companies out there will add a non-compete clause.
- dragoncrab 3y agoNot in the US or EU. However, they can still sue you in the US and due to the broken legal system there, it will cost you decent money even if they are bound to lose from day 1.
- wkat4242 3y agoIn Holland too. Even if someone sues you maliciously, you still have to pay the state a fee to be heard. Otherwise the judgement will fall to the enemy party by default. You can recuperate this cost from them when you win but you're still out of pocket for your time and the money until you manage to cash it from them which can be hard. And they can keep doing it. The system is very unfairly biased in favour of people with lots of money
- Animats 3y agoAs someone who's into virtual worlds, and a user of Second Life, it's impressive to see how well those systems stay up. There hasn't been a total outage of Second Life in 5-10 years. Once Amazon's networking went down in a way that prevented new logins for a whole day, but existing logins remained. The 3D world, which has a lot of stuff going on even with no users around, continued to work. This is an extremely complex one of a kind system, and it just keeps cranking along. It's very distributed; one region (a 256x256m square) can crash and restart without taking down its neighbors. Users see the failed region as a square hole in the ground filled with water until the server restarts, which takes about two minutes. So outages are quite graceful. It's currently hosted on AWS, but it doesn't have to be. What fails? The associated webcrap. The Marketplace, which is just a catalog and shopping cart. The forum system, which is outsourced to Invision, seems to fail several times a month. The messaging system, which is just a lightweight social network. The billing system. The outgoing payments system. Amazon's outgoing HTTPS proxy. All of those have failed several times in the last year. Even the JIRA system conked out once. The quality of web software is underwhelming.
- supriyo-biswas 3y ago> Amazon's outgoing HTTPS proxy. Is this ELB or Cloudfront we’re talking about?
- phartenfeller 3y ago> The quality of web software is underwhelming I hate where "web scale" has brought us. Because some 0.01% of giants have tremendous scalability problems, every small project needs an overly complicated architecture that consists of layers of services. In the end, nobody understands the monster, and the complexity brings more issues than it solves. But still, this is somehow a standard today. I have a lot of love for static HTML sites or simple backend/frontend solutions. The web is great, but current development trends are not.
- havkom 3y agoYep. If you keep off the “latest trends frameworks” and just keep it as vanilla and simple as possible web development can be productive, scalable and pleasant.
- deleted 3y ago[deleted]
- hermitcrab 3y agoWe are a small software company (2 people) and we've also had plenty of issues with Google over the years. Mostly related to Google Adwords. For example: https://successfulsoftware.net/2015/03/04/google-bans-hyperlinks/ https://successfulsoftware.net/2015/03/04/google-bans-hyperl... https://successfulsoftware.net/2016/12/05/google-cpa-bidding-goes-wild/ https://successfulsoftware.net/2016/12/05/google-cpa-bidding... https://successfulsoftware.net/2020/08/21/google-ads-can-charge-you-anything-they-like-for-a-click-on-their-partner-network/ https://successfulsoftware.net/2020/08/21/google-ads-can-cha... https://successfulsoftware.net/2021/05/04/wtf-google-ads/ https://successfulsoftware.net/2021/05/04/wtf-google-ads/ If Google have no interest in providing decent support to the author of the original article, who are paying megabucks to Google, what hope do small businesses like mine have?
- biorach 3y ago> Google have no interest in providing support
- hermitcrab 3y agoNo, they want to do things on a massive scale. Which means it is really difficult to talk to a human for support. And even if you manage it, it might be some badly trained subcontractor. But somehow there are always humans available to ring you up and tell you how you can spend more on Google Adwords.
- strstr 3y agoSounds like a genuinely frustrating experience. Bit confused about why nested virt has anything to do with their problems given that they aren’t using virt inside the VMs. Softlocks are a generic indication of a lack of forward progress. Same confusion with the MMIO instructions comment. If that’s about instruction emulation, not sure why it matters where it happens? It’s both slow and bound for userspace anyway. If it’s supposed to be fast it should basically never be exiting the guest, let alone be emulated. Sounds like the author is a bit frustrated and (understandably) grasping at whatever straws they can for that most recent incident.
- supermatt 3y agoNo doubt all cloud providers have their problems. For my day job, over the last 2 years we have discovered and reported multiple issues with Keyspaces, Amazon Aurora, and App Runner. In all cases these issues have resulted in performance degradation, and AWS support wasting our time sending us chasing our tails. After many weeks of escalation, we eventually ended up with project leads who confirmed the issues (some of which they were already aware of, yet the support teams had wasted our time anyway!) and (some of them) have since been resolved. We are stuck with Keyspaces for the time being, but now refuse to use any non core services (EC2, EBS, S3). As soon as you venture away from those there be dragons.
- wavemode 3y agoOh, for goddamn sure. Half the services on AWS, probably, are very poorly designed or very poorly run (or both). CloudWatch stands out to me as one that is mind-bogglingly buggy and slow. To the point of basically being a "newbie trap" - when I see companies using it for all their logging, I assume it's due to inexperience with the many alternatives. At least the compute services are reliable.
- yibers 3y agoI actually use cloudwatch quite a lot. I didn't notice many bugs or slowness, but I assume I am missing something. Can you perhaps point to some specific issues you had with Cloudwatch?
- rubiquity 3y ago$$$ and a strange API. The internal metrics service that Amazon has used for ages works much better for power users. CW is slowly becoming like it.
- wavemode 3y agoThe user interface is unintuitive, text search is slow, querying is even slower, refining the metric graphs to a time period is really annoying, the graph controls are consistently buggy in my experience... those are just off the top of my head
- wg0 3y ago> In 2022, we experienced continual networking blips from Google’s cloud products. After escalating to Google on multiple occasions, we got frustrated. So we built our own networking stack — a resilient eBPF/IPv6 Wireguard network that now powers all our deployments. Suddenly, no more networking issues. My understanding is that the network is a VLAN programed via switches for VMs so when you create VPC, you're creating a VLAN probably. So how can an overlay (UDP/Wire guard) be more reliable if the underlaying network isn't stable? PS: Had even 1/10th of issues have happened on AWS with such a customer, their army of solution architects would be camping in conference rooms every other week reviewing architecture, taking support engineers on call and what not.
- devsda 3y agoMy guess is that whatever clever network optimizations that Google has are probably interfering with their traffic. By building their own network stack, they are skipping them and also wireguard might be better equipped to dealt with occasional faults as it built on udp which is inherently unreliable.
- Bluecobra 3y agoI have a direct cross connection to Google in a colocation facility (aka Dedicated Interconnect). One issue I found is that Google would randomly shuffle around their BGP routers which would cause BGP to flap and briefly losing all connectivity. When I raised this issue with support their answer was that this is expected behavior and we need to purchase a redundant connection. Mind you this isn’t cheap, we’re talking around $2,000 per month for a 10G connection when you add up all the GCP/colo fees. It’s pretty laughable that they can’t preserve TCP connections when they migrate their cloud routers around. I have had BGP uptimes on direct cross connects for over a year with other vendors on bare metal.
- bushbaba 3y agoIt’s about scale. Google was built for an order of magnitude greater scale where such reliability of a single link would be cost prohibitive. However in general if you need high uptime, you’ll need multiple peering links. AWS and azure also recommend the same.
- simo7 3y agoInteresting, I’m starting to think undocumented thresholds are quite common in GCP. I experienced something similar with Clod Run: inexplicable scaling events based on CPU utilization and concurrent requests (the two metrics that regulate scaling according to their docs). After a lot of back and forth with their (premium) support it turns out there are additional criteria, smthg related to request duration, but of course nobody was able to explain in details.
- politelemon 3y agoUnnanounced changes too, there was a Firefox outage in 2022 due to GCP: https://hacks.mozilla.org/2022/02/retrospective-and-technical-details-on-the-recent-firefox-outage/ https://hacks.mozilla.org/2022/02/retrospective-and-technica...
- merb 3y agosorry but the blame here was 100% on Mozilla. No matter which http version, headers should always be treated as case-insensitive. Blanking anything on google here is just stupid. The problem was nih-syndrome and ignored the http spec.
- mst 3y agoMozilla are entirely clear that this was their bug. However, GCP changing the default under their infrastructure without prior warning was still unacceptable. Operations work should (IMO must) be conducted with the expectation that any major change like that will expose existing bugs in deployed code. (I've done enough ops work in my life that I'd love to say 'will potentially expose' but in practice there's always -something- that breaks and if I don't find it in the first 24h after a major change I'm going to spend the next two weeks waiting for the shoe drop to happen)
- merb 3y agoGCP does send mails when you abo‘d them. GCP is not to blame if they used auto. Heck if your loadbalancer sends you headers lowercase with a new http version it should not result in a bug. GCP‘s change was fine. Their software had a bug that would‘ve led to request smuggling.
- StopHammoTime 3y agoI have a lot of interaction with Google Cloud Support, mostly around their managed services. I am genuinely not over-impressed with their service, considering with similar employers of size on AWS the support experience was always wonderful. However, I will say if you are on Google Cloud and you have a positive interaction, make a big deal about someone helping you. Given the rarity it occurs, it’s not a big deal to really go out of your way to reward someone with some emphatic positive feedback. I’ve had four genuinely fantastic experiences and there’s always a message to a TAM that flows soon after. I hope more people like those I interacted with get rewarded and promoted.
- latchkey 3y ago> However, I will say if you are on Google Cloud and you have a positive interaction, make a big deal about someone helping you. This. These sorts of discussions are like bike shedding over vi/emacs. Only the complaints make it to the front page on HN. I've been using GCP off and on for projects for a decade now. Built multiple very successful businesses on it. Sure it hasn't been all perfect, but I'm an overall happy camper. Having also used AWS heavily when I was on the team building the original hosted version of Cloud Foundry, I'd never go back to them again. It was endless drama.
- Kwpolska 3y agoYou should've migrated many months ago, if a cloud provider forces you to build your own networking or registry, you shouldn't use that cloud provider.
- politelemon 3y agoThat was the first thing that struck me, the 'workarounds' stagger belief, but they seem to be casually dropped in (?). If I were in a situation where my company was contemplating implementing building our own registry/network stack, then the benefits of using a cloud provider are gone, and I would have considered moving to another provider... not saying "I can fix him". This feels like a sunken cost perhaps that is the right term.
- mrj 3y agoI would bet that they were thinking about colocating and would need to have secured intra-service communication anyhow. In Google this is transparent but at a (probably yet identified facility) it'd be up to them to provide.
- justjake 3y ago(Blogpost Author) Yup. We've been thinking about colocation for a while, so we've just been building these up. Basically all that's left is to make our volume storage bulletproof. We'll do that as we're moving stateless workloads to bare metal early next year, and ideally be off GCP EOY latest
- supriyo-biswas 3y agoWell for folks building out cloud infrastructure, building your own networking stack and registry is a good way to achieve platform independence, without which you'll be left at a disadvantage and vulnerable to the whims of cloud providers who may or may not extend volume discounts, thus indirectly harming your ability to compete.
- chrisandchris 3y ago
- kgeist 3y ago>We have automated systems in place to detect and resolve this. We’re notified in Discord Isn't Discord hosted on GCP, too? If it goes down, monitoring also goes down?
- justjake 3y ago(Blogpost Author). We use Discord to notify. Our monitoring runs directly to PagerDuty for anything we actually need to action on.
- rurban 3y ago> In our experience, Google isn’t the place for reliable cloud compute, and it’s sure as heck not the place for reliable customer support. Always was, always will be. For them customers are always the last
- asylteltine 3y agoGCP is the ugly stepchild of Google. They prioritize their own infra which unlike Aws doesn’t even run on gcp! It’s a joke. They don’t dogfood anything. All Google infra runs on separate systems (both) or dedicated deployment like their own spanner clusters. Google employees look down on gcp employees like second class citizens
- mst 3y ago> They prioritize their own infra which unlike Aws doesn’t even run on gcp! I don't believe that to be the case, last I heard a year or two back the vast majority of it -does- run on a google GCP tenant account and what didn't was largely at least in the process of migration planning. (my source here is "pillow talk with a senior GCP engineer" and I don't believe she had any reason to lie to me)
- asylteltine 3y agoThat’s not been my experience having done consulting for Google. They may have some stuff on gcp but they don’t dog food much like Aws does (literally running Amazon on top of vanilla dynamo and kinesis). Google has a lot of custom infra
- anonacct37 3y agoSorry, but that's not true. I wish it was. But running anything that bridges google3 and GCP is a nightmare. Outside of acquisitions and OSSish stuff like chrome, it's really rare to see GCP used. Oddly enough their corp eng team does quite a bit with making GCP accessible to the rest of the company. Source: former Google SRE who actually worked on one of those teams.
- Sytten 3y agoWhishing all the best to the railway team, they really are building something nice. Hopefully the move to bare metal will mean price reductions for customers. I am philosophically opposed to cloud providers charging per user on top of very expensive resources but it might just be me.
- nomilk 3y ago> In our experience, Google isn’t the place for reliable cloud compute In the early days of cloud computing unreliability was understandable, but for Google to be frustrating its large customers in 2023 is a pretty bad look. Curious to know if others have had similar experiences, or if the author was simply unlucky?
- whirlwin 3y agoI don't know how it happened, but I used GKE for a side project, which was overkill for such a small project, and I could live with $100 /month, but the bill kept creeping up to $300 and later $400 with no apparent explanation or workload increase. I had no choice but to revert to something else. Ended up with good old Heroku with $20/month and never regretted it
- deleted 3y ago[deleted]
- lawgimenez 3y agoIf you go to Google’s issue tracker, you will find a lot of issues that were ignored. For example, this [0]issue that caused our ANR rate to dip. [0] https://issuetracker.google.com/issues/230950647 https://issuetracker.google.com/issues/230950647
- doubloon 3y ago"reasons why Oxide has a business #12390"
- asylteltine 3y agoAn oxide rack has a minimum cost of something like 600k not including all the infra you need to run a rack, maintenance, and then needing to upgrade
- mst 3y agoRailway's bill was into the multiple millions per year at the very least so that doesn't necessarily rule it out.
- asylteltine 3y agoThat’s one misconception about leaving cloud people think it’s a one time cost compared to opex, but in reality you are just moving the spending. Devops for your now custom on prem workflows, toil due to inferior tooling compared to cloud, physical costs like electricity and cooling, space for the racks, physical security, high availability, etc I mean there’s so much downside.
- mst 3y agoI said "doesn't necessarily rule it out" rather than a stronger claim advisedly. You're entirely correct that there are an unfortunate number of people who hold that misconception, but (a) I'm not one of them (b) that wasn't my point.
- mbStavola 3y agoIn the post they say they pay Google "multiple millions" of dollars already. Depending on their needs, the TCO of Oxide racks may end up being less than what they pay GCP.
- latchkey 3y agoThat's just moving the goal posts around.
- fidotron 3y agoMaybe it is me but this doesn’t exactly reflect well on anyone. Isn’t the value prop of railway not having to worry about things like this? It doesn’t matter what the problem is - you shouldn’t be passing such problems on to customers at all. I have worked on a product that caused such a spike on Google App Engine that within 20 minutes of it going public Google were on the phone explaining their pagers all went off, and in that case resolved to temporarily bump the quota up for 48 hours while a mutual workaround was implemented. The state of Google Cloud today seems just another classic case of the trend of blaming the customer.
- davidgerard 3y agoHOW TO CHOOSE A CLOUD PROVIDER * AWS: you will pay to have stuff work properly and you like having customer service * Azure: you hate yourself, you're running Windows or both * Google: you're cheap enough that basic functionality is an optional extra * Oracle: lol * Hetzner: cheap, good service, the finest pets in the world, no cattle
- diamondfist25 3y agoWhat about DO?
- monlockandkey 3y agoYou should always reach out to use Digital Ocean, Linode, Vultr as your starting point. Aws and the gang are mega mega expensive compared to what you pay for a vps. If you require services beyond compute, database and storage, then use the big names. Otherwise save yourself headache with complexity, unpredictable and absurd costs. Please don't use AWS especially as a startup, you are going to kill yourself paying for compute, database and egress that is multiples times what you get from a vps. AWs is NOT cheap
- davidgerard 3y agomy personal site is on Hetzner, fwiw. they are extremely good IME
- monlockandkey 3y agoYes Hetzner is a excellent choice
- nijave 3y ago>Azure Or you have big enterprise customers that have a grudge against Amazon and Google and refuse to use anything else.
- 3y ago
- ghusto 3y agoI know AWS isn't cool or sexy, but shit works.
- motoboi 3y agoIt’s sad to see people rediscovering that GCP is not a serious product over and over again.
- markbnj 3y agoWhat's your personal experience with it? We've been on the platform for almost eight years. Three clusters, hundreds of compute VMs, four or five public and private DNS zones, 10+ cloud sql dbs, about the same number of memorystore and firebase instances. We just don't see these issues as related in the OP, and when we do have problems support has been fast and helpful. Not to gainsay their experience, since it was obviously frustrating, but truthfully you can find similar stories about all compute providers.
- motoboi 3y agoNot everyone get cancer from cigarettes you know? See the avalanche of horror stories. I had mine.
- acdha 3y agoThat’s what I was thinking. Adopting GCP after years on AWS was eye-opening: I’d previously had a good impression but it was a constant cycle of “oh, we don’t have that - build your own” and issues which have been open for years full of customers asking and GCP PMs stalling. Then the price increases started, making it even harder to defend paying more for less.
- asylteltine 3y agoI work at a company that spends billions on AWS and we intentionally have minimal gcp deployments and ban compute there because of how unreliable gcp is and how awful (outsourced) their support is. Gcp has excellent products but garbage operations. Who is running that clown show? It could have easily been the #2 cloud if they knew what they were doing
- ransom1538 3y ago"On December 1st, at 8:52am PST, a box dropped offline; inaccessible. And then, instead of automatically coming back after failover — it didn’t. Our primary on-call engineer was alerted for this and dug in. While digging in, another box fell offline and didn’t come back" This makes no sense. A machine restarted and you had catastrophic failure? VMs reboot time to time. But if you design your setup to completely destroy itself in this scenario, I don't think you will like a move to AWS, or god forbid, your own colo.
- wavemode 3y agoRead the article more carefully. The article (the text you quoted, even) clearly states that the machine didn't "restart". It crashed and didn't come back online. And nowhere in the article do they state that this was a "catastrophic failure" - Railway itself didn't go down entirely. But Railway is a deployment company, so they are re-selling these compute resources to their customers to deploy applications. So when one of those VMs goes down and doesn't automatically failover, that's downtime for the specific customer who was running their service on that machine. As they state: > During manual failover of these machines, there was a 10 minute per host downtime. However, as many people are running multi-service workloads, this downtime can be multiplied many times as boxes subsequently went offline. > For all of our users, we’re deeply sorry.
- xyzzy_plugh 3y agoTFA is a bit too light on details. Boxes due, it's a fact of life. I don't really follow what "didn't come back online" is supposed to mean. Nodes aren't 100% durable. A lot depends on your particular configuration. In any case, there's no world where all VM failures trigger automatic reboots. Expecting that to be the case just makes no sense. Automatically failing over should be handled at another layer, for which there a many possibilities. Manually restoring nodes sounds like a "pets, not cattle" problem. Long ago, we used to run into this on AWS all the time before we started automatically aging them out.
- deleted 3y ago[deleted]
- annoyed_eng 3y agoGenerally, I think over the last few years, GCP has lost its way. There was a time several years ago where they were a meaningfully better option when looking at price / performance for compute / storage / bandwidth when compare to AWS. At the time, we did detailed performance testing and cost modeling to prove this for our workload (hundreds of compute engine instances etc). Support back then was also excellent. One of our early tickets was an obscure networking issue. The request was quickly escalated then passed from engineers in different regions around the world until it was resolved. We were very impressed. It was a change on the GCP end that ended up being reverted. We quickly got to real engineers who competently worked the problem with us to resolution. The sales team interactions were also better back then. We had a great sales rep who would quickly connect us with any internal resources we needed. The sales rep was a net positive and made our experience with GCP better. Since then, AWS has certainly caught up and is every bit as good from a cost / performance standpoint. They remain years ahead on many managed services. The GCP support experience has degraded significantly at this point. Most cases seem to go to outsourced providers who don’t seem able to see any data about the actual underlying GCP infrastructure. We too have detected networking issues that GCP does not acknowledge. The support folks we are dealing with don’t seem to have any greater visibility than we do. It’s pathetic and deeply frustrating. I’m sure it’s just as frustrating for them. The sales experience is also significantly worse. Our current rep is a significant net negative. We’ve made significant investments in GCP and we hate seeing this happen. While we would love to see things improve, we don’t see any signs of that actually happening. We are actively working to reduce our GCP spend. A few years ago, I was a vocal GCP advocate. At this point, I’d have a hard time suggesting anyone build anything new on GCP.
- vel0city 3y agoIt's hilarious people are bashing GCP for having one compute instance go down and the author acknowledges it's a rare event. On AWS I've got instances getting forced stopped or even straight disappearing all the time. 99.95% durability vs 99.999% is way different. If they had the same architecture on AWS it would go down all the time IME. AWS primitives are way less reliable than GCP, according to AWS' docs and my own experiences.
- Wuzado 3y agoThe article doesn't seem to mention AWS, really. I also feel like the primary issue is the lack of communication and support, even for a large corporate partner. Seems like they're moving to bare-metal, which has an obvious benefit of being able to tell your on-call engineer to fix the issue or die trying.
- vel0city 3y agoBut in this case the answer from AWS would have been that's their SLA and you need to just be ready to handle an instance getting messed up from time to time, because it's guaranteed to happen.
- deanCommie 3y agoEC2 [0] and GCP Compute [1] have the exact same SLAs, which is 99.99%, dipping below which gets you a 10% refund. Dipping below 95% gets you a 100% refund. [0] https://aws.amazon.com/compute/sla/ https://aws.amazon.com/compute/sla/ [1] https://cloud.google.com/compute/sla https://cloud.google.com/compute/sla
- vel0city 3y agoBy the links you shared instance level SLA on AWS is 99.5%. GCP instance level is 99.99%. That's not the same. > For each individual Amazon EC2 instance (“Single EC2 Instance”), AWS will use commercially reasonable efforts to make the Single EC2 Instance available with an Instance-Level Uptime Percentage of at least 99.5% The underlying storage isn't the same as well, and that matters more. EBS is 99.95% durable. Even standard zonal PD's on GCP are >99.99%, balanced are >99.999%, SSDs are >99.9999%. Even if it was 99.99% (it's not on AWS) what's the point of having your instance be 99.99% if the underlying disks might disappear? That's something I've seen happen multiple times on AWS, never once on GCP.
- tlogan 3y agoAll these cloud service providers have bugs and issues. But the problem with Google is that their support seems somehow disconnected from the real world. There is support, and they do respond to chats, calls, or emails. However, it often feels like I'm talking to someone who doesn't genuinely care about my concerns or do understand what I’m talking about. Good support is hard to come by and hard to implement. So I really don't know what is missing in Google's support that exists in AWS support. Maybe because AWS support staff are trained to first put themselves in the customer's shoes and understand the problem from my perspective.
- deleted 3y ago[deleted]
- M_bara 3y agoThe mantra at aws is customer experience. Any time there’s an outage or impact, the first numbers to be stated are customer impact related. In fact, as an engineer you might be empowered enough to recommend a refund for a customer (even though the customer may be at fault) and the refund will go through. Full disclosure: worked for aws devops for a couple of years
- niuzeta 3y agoI wonder how many of these stories it would take before it starts affecting Google's bottom line. I've tinkered with GCP on small side projects, sure - but after exposure of these stories for over a decade in HN, I can never recommend GCP as a serious cloud alternative. I can't imagine I'm the only one in this boat.
- tedd4u 3y agoIt sounds like if you deploy on Railway they don't automatically handle a box dying (e.g. with K8s or other) -- "half the company was called in to go through runbooks." When they move to their own hardware, how will they handle that?
- londons_explore 3y agoGCP is pretty reliable - for a smallish deployment you could probably go a couple of years before seeing a machine die. So they probably never built in health checks and auto fail over.
- testernews 3y ago“ We paid them multiple millions of dollars per year” Never heard of railway but paying this many $$$ per year should give you a dedicated support rep. But google doesn’t do support for anything lol
- londons_explore 3y agoOh - they give you a support rep. Just the support rep is powerless to do anything.