9 ms·
Oracle Cloud is having a major outage
What's worse is that their "real time" status is not showing any issues.
https://ocistatus.oraclecloud.com/#/
I had to confirm the outage based on community reported down detector.
https://downdetector.com/status/oracle-cloud/
All of our services, instances and backups for https://searchadsoptimization.com are in Oracle cloud.
This shows a critical issue when relying on a single Cloud provider. It's time to build a cross cloud infrastructure design to handle these issues.
Update: It looks like they have updated their “real time” status page after good 25 minutes of severe outage. My trust and assumptions with real time status pages changed completely.
I don't understand the point of real time status pages if they are clearly not real time and not accurate.
My error notifications were blowing up my phone, the first thing I did is check their status page and assumed issue is within my application, and I couldn't even access my backend application. Out of desperation, I had to check downdetector to confirm the issue. I have formed new respect for downdetector.
- berbec 3y agoIt's now showing errors in US East (Ashburn)
- internetter 3y agoWell, if the status page says they are up do the SLAs really apply? (half kidding, this seems like an oracle thing to do)
- resdev 3y agoI don't understand the point of real time status pages if they are clearly not real time and not accurate. My error notifications were blowing up my phone, the first thing I did is check their status page and assumed issue is within my application, and I couldn't even access my backend application. Out of desperation, I had to check downdetector to confirm the issue. I have formed new respect for downdetector.
- gtirloni 3y ago1. It only goes to red after a set of humans determine it's really high impact and should be made public. Minor or localized outages rarely qualify. 2. Previous point is ignored very often and outage is only made public when major clients or news organizations take notice and inquire.
- Spivak 3y agoI like AWS's approach of having your own personal incidents page. Still not exactly real time but better than an unchanging wall of green. And they include performance degradations as incidents which is nice.
- ben7799 3y agoAs a former employee this REALLY seems like something Oracle would actually do, along with instructing employees to tell customers it wasn't actually down. It was before Oracle cloud but I literally was told to do things like that.
- deleted 3y ago[deleted]
- firstSpeaker 3y agoThey need SVP+ level approval to mark disruption/unavailability.
- stusmall 3y agoI love this because I honestly can't tell if its a joke. It's obviously a terrible idea.... but also seems like something Oracle would do.
- tatersolid 3y agoAWS does the same; it stays green unless the outage is so bad an exec has to approve changing the status page. Surprisingly azure is very open with outages of all services big and small in my experience, and notifies if any service our tenant is using is impacted.
- jareds 3y agoNot all applications need high availability across multiple clouds and the cost increases that go with that. Some applications can afford a couple hours of downtime if the underlying hosting platform as issues instead of needing to do a hot failover to completely different infrastructure.
- resdev 3y agoI agree, I should've been more clear, i was referring to database. The main issue is if database , storage and backups are located with single Cloud provider, it is a possibility for a single point of failure.
- cheeze 3y agoThe idea that "we'll just seamlessly failover to another provider" is a bit otpimistic IMO. With that comes additional complexity. Some applications need this, but it's a huge cost and complexity tradeoff for almost all businesses. I'm a fan of sticking with one provider, but going with something bigger that has a good track record. AWS, GCP, Azure aren't prone to 0 outages, but I think for almost all companies, having redundant stacks in separate regions is enough to maintain high availability. I don't know enough about Oracle Cloud to comment on them, but my general take is these companies all inevitably hit a "showstopper" global outage, realize they aren't investing enough in separation of regional stacks enough, and put a ton of energy into making their platforms more fault tolerant. Thinking that Johnny dev shop is going to be able to do better than a major player is, IMO, wishful thinking. I know that at GCP at least, they actually have monitoring setup for things like tweets, downdetector, etc. Ideally they catch every issue with their own monitoring, but they do their best to know if anyone is having an issue, whether they can detect it or not..
- gtirloni 3y agoCross cloud is really complex and error prone. You'll probably cause more outages than you'll prevent by going down that path. Maybe you should consider moving to a major cloud provider that has better services.
- AnthonyMouse 3y agoIt's not really that bad? If you're operating at large scale you should have enough control over your own infrastructure to distribute load to multiple providers. Then if one of them is down, spin up more instances on another one. If you're operating at small scale, store your backups on another provider, and periodically test that you can quickly restore them to that provider. This isn't just about redundancy. Doing this is necessary to keep you from getting locked in.
- gtirloni 3y ago> If you're operating at small scale, store your backups on another provider, and periodically test that you can quickly restore them to that provider. Sure, but that's not what people mean when they say cross cloud. It usually implies running active workloads. Orchestrate hundreds of workloads that each depend on one another across several clouds and they reorganize upon a failure introduces a lot of new failure modes.
- AnthonyMouse 3y ago> Sure, but that's not what people mean when they say cross cloud. It usually implies running active workloads. For workloads large enough to justify load balancing, that isn't that hard. Things should only rarely care which "cloud" they're running on. The amount of the load assigned to each one is a knob you can turn. For workloads smaller than that, the primary way to achieve redundancy is failover. Whether it's worth the cost and complexity of making that happen automatically instead of manually doesn't have a universal answer. But "just use a more reliable provider" doesn't work if being down matters. AWS, Azure and Google Cloud have all had major outages. More than that, sometimes a piece of equipment fails in a way that takes adjacent equipment with it, or there is a fire or a burst pipe. They call it "cloud" but somewhere there is physical hardware under your bits and it can fail. Each failure may only affect a limited number of customers, but those customers can include you. If your systems can't be down, you need a plan in place to have them running again somewhere else in short order. And putting the system you use for this on a different provider can save you from a major provider outage.
- hn_throwaway_99 3y agoWow, there must be literally tens of people who are worried right now! My shitty, snarky comment aside, I am genuinely curious about why someone would choose Oracle as a cloud provider. If you look at their capex spend, it's undeniable they have so vastly underinvested in their cloud compared to AWS, Azure and GCP, that even if you were an "Oracle shop" I'm genuinely curious what benefits their cloud would offer. Edit: Just want to say I really do appreciate the responses, lots of good info! I didn't know Oracle cloud offered a decent free tier, will take a look.
- remram 3y agoThey have a good free tier: ARM Ampere instance with 24 GB memory.
- natrys 3y agoEven in pay as you go they have it much cheaper than in azure or gcp.
- space_ghost 3y agoI snagged one of those as soon as they were available. And, so far, I've only used it to host my portfolio, which is entirely static. :D
- yabones 3y agoI abuse mine to run a big-ass Elasticsearch instance with about a year of syslogs/weblogs for my personal machines. Not terribly useful, but a good outlet for my hoarding I suppose.
- BryantD 3y agoHosted Oracle database services. Otherwise, very favorable contract terms.
- hn_throwaway_99 3y agoThanks very much, was assuming they must be giving a sweetheart deal given that AWS and GCP have hosted Oracle DB solutions, and even Oracle itself touts running on Azure, https://www.oracle.com/cloud/azure/oracle-database-for-azure/ https://www.oracle.com/cloud/azure/oracle-database-for-azure....
- deleted 3y ago[deleted]
- not_enoch_wise 3y agoIf you can’t trust Oracle, who can you trust?
- hn_throwaway_99 3y agoEven people/companies who use and depend on Oracle, I wouldn't quite say their relationship is one of "trust".
- dralley 3y agoI'm pretty sure that's the joke.
- compumike 3y ago> My trust and assumptions with real time status pages changed completely. FYI this is why we show real-time status on https://heiioncall.com/status https://heiioncall.com/status including the time of the last inbound check-in or last HTTP probe.
- internetter 3y agoThis is a beautiful product with excellent pricing and the design is lovely. Great work.
- s-xyz 3y agoVery honest question, who would use Oracle Cloud in 2023?
- hamburglar 3y agohttps://downdetector.com/status/tiktok/ https://downdetector.com/status/tiktok/
- phendrenad2 3y agoMaybe Larry plugged in his solar-powered yacht backwards and took out the local grid.
- qwertyuiop_ 3y agoOracle Cloud Infrastructure Customer, We've identified a cooling system issue affecting multiple services in the US East (Ashburn) region. Our engineers are actively working to mitigate the issue.
- sicklife 3y ago......
- TX81Z 3y ago[dead]
- hu3 3y agoI have a VM there but it wasn't affected. Got this e-mail from Oracle 50 minutes ago: > Oracle Cloud Infrastructure Customer, > Engineers and the colocation partner have successfully installed additional cooling systems to reduce ambient temperatures and mitigate the issue affecting multiple Oracle Cloud Infrastructure (OCI) services in the US East (Ashburn) region. We will continue to closely monitor this situation.
- bastard_op 3y agoThe reason for outage report should be interesting. My cousin mentioned their erp was down mid-day, and I laughed citing HN like "oh yeah, forgot you're a poor bastard oracle user." It was entirely dead, like everything apparently, most of the day. Sadly the financial people don't care, they will still cut a check to daddy Ellison monthly. At least one large California municipality I worked with made a multi-year concerted effort to abandon the misery that is oracle erp. That said, never heard how that venture panned out with the replacement. Something about a frying pan to the fire comes to mind.
- bastard_op 3y agoSo did any actual RFO come out about this yet? Inquiring minds want to know, also point and laugh.