11 ms·
IBM Cloud was down, as well as their status page
- Lyren 6y agoI received communication ~15min ago that they're actively looking into the issue. I submitted the ticket roughly 20min ago. So it seems they're aware. It doesn't help that their status page is also hosted on IBM Cloud.
- woakas 6y agoOur site (ubidots.com) does not have a complete down, but the IBM network has a high latency.
- cerw 6y agoBeen like that for last 1h, Network packet Sydney (GCP) to Sydney (IBM) 62% packet loss
- leetrout 6y agoHugops. Hope they get a root cause and a quick fix. I’m not a fan of their cloud service but I know people working on the outage and fix are stressed.
- voz_ 6y agoI generally do everything on AWS or GCP, with a little Azure sometimes for personal projects. In what world does IBM beat one of those three in anything? Generally curious - how they are able to stay competitive?
- wmf 6y agoThey had bare metal before Packet or AWS and inter-region traffic is free.
- Operyl 6y agoTheir biggest thing going for them is 100% free dark fiber private network. You do have to pay for the bigger pipe (100mbps included for each server, a minimal upcharge for gigabit), but that's pretty much a rounding error.
- blantonl 6y agoSoftlayer
- twalla 6y agoTheir bare metal cloud offering (SoftLayer acquisition) was actually pretty good whenever I used it about 4 years ago. Wasn’t the most intuitive API or UI but you could get a bare metal server anywhere in the world in a few minutes.
- rad_gruchalski 6y agoWhen the wind blows in the right direction. Sometimes, your server would get stuck in provisioning for hours and only get „un-stuck” after creating a support ticket. Which, I kid you not, at one of the previous jobs, wd had automated in our provisioning popeline. Good times. But when it worked, it worked. API was voodoo.
- Kinrany 6y agoFixing provisioning based on support tickets might have been automated on their side too :)
- bashinator 6y agoI just discovered this today: aws support create-case \ --subject "not working" \ --communication-body file://description.txt
- jonfw 6y agoThey've got the only real managed Openshift option right now, and their managed Kubernetes services is really great and seamless IMO.
- colinbartlett 6y agoMy side project StatusGator monitors status pages (including IBM's ill-fated page) and I'm seeing more than 10% of the nearly 800 services we monitor having an outage right now. So it appears to affect anyone who depends on IBM Cloud.
- Jedd 6y agoBig fans of StatusGator here. Do you have similar %'s of monitored cloud services that have gone off the air during other providers' outages?
- colinbartlett 6y agoThanks and that's a GREAT idea for some detailed analysis. I have been trying to make better use of the 5+ years of status page history stashed in the cupboard.
- bigiain 6y agoThis is something I'd totally read. (You should headhunt the guy at BackBlaze who does their hard drive stats blog posts, and release this data analysis quarterly!)
- atYevP 6y agoYev from Backblaze here -> please do not do this, we really like him.
- brian_herman__ 6y agoYeah I think the backblaze does great work for backblaze for their hard drive report I can’t wait for him to talk about the WD SMR fiasco!
- bigiain 6y agoHow about: (You should contact the guy at BackBlaze who does their hard drive stats blog posts, and pay him to do this for you as a side hustle - and release this data analysis quarterly!) :-)
- viejoqueso 6y agoMaybe they should also shut down their cloud.
- julianeon 6y agoSeems pretty dumb to host a status page in a way that it could go down, when it should be a static page that is trivially hosted on CDN's worldwide.
- koolba 6y agoYou can’t cache it for that long though. A better approach is to have it hosted on a different cloud platform. If you really care, you’ll set it up on a different domain and nameserver as well with a long lived redirect (cached on CDNs) from the usual status.example.com or example.com/status.
- julianeon 6y agoThanks; you're right - the caching would be a problem, so your solution makes more sense.
- syshum 6y agoOver confident in their own Cloud "Our cloud can never go completely down We are IBM, we have Watson..."
- sky_rw 6y agoTheir status page also seems integrated into their internal support ticketing system. It's not a traditional status page. They wanted to maintain a consistent garbage interface to keep it inline with the rest of their administrative service.
- blantonl 6y agoAll of Broadcastify's audio servers (hosted with Softlayer in their Dallas datacenter) are completely unreachable and down. I'm going to wait a bit to see if we get a status update, otherwise we'll be spinning up instances on AWS to failover (which will be enormously costly for bandwidth) No status, no nothing, we're in the dark.
- Operyl 6y agoHey. Do you want to shoot me an email, IRC chat, or anything? I can keep you up to date with what I'm hearing from my manager.
- dashesyan 6y agoHey, I'm a customer of IBM Cloud, too. Could you share what you're hearing from them? It would be nice to know what's going on
- Operyl 6y agoSo far? Pretty much no news, they're using Slack to communicate a bit. VPN access for everybody is broken or barely working, no internal ticketing as a result. As far as I can tell, private networking is mostly working between servers (at least, for my servers, ~60).
- chadcmulligan 6y agoAfter signing the appropriate NDA's and being vetted by IBM legal of course
- Operyl 6y agoEverything I'd say would be public knowledge, of course. I don't work for IBM, I'm a user :).
- chadcmulligan 6y agoYou never know which way levity will go on Hacker news :-)
- AaronFriel 6y agoAh, is this the exception that proves the rule that "no one was ever fired for buying IBM?" Sorry to be glib, I'm sure it's a tough time for people who were sold on their cloud platform and work on it!
- mark-r 6y agoEverybody's cloud goes down sometime. The big fail here was hosting their status page on the same infrastructure.
- oceanswave 6y agoBut usually only a single AZ or region... seems like this is bigger?
- thephyber 6y agoHow sure are we that this outage is limited to IBM cloud? Pindom[1] had a spike of website outages from 11k => 27k. [1] https://livemap.pingdom.com/ https://livemap.pingdom.com/
- Nextgrid 6y agoIt's most likely customers of IBM cloud whose systems rely on something hosted there and are thus down as well.
- thephyber 6y agoYes, I considered that possibility before posting.
- rat9988 6y agoI'm not sure what you are trying to prove with your comment then.
- TallGuyShort 6y agoSometimes people ask questions when they aren't sure and don't have anything to prove.
- deleted 6y ago[deleted]
- thephyber 6y agoMy top-level comment was trying to ask for any evidence whether the downtime was related to a wider network issue (eg. BGP, backbone) or if it was specific to (part of) the IBM cloud. The link was just a data point, not evidence of anything in particular.
- caiobegotti 6y agoHonest slightly cynical question: most probably someone inside the responsible team said some day that it would be very stupid to host the status page inside the same infrastructure being monitored, but they were probably ignored... what should that person do now? Say "toldya!" out loud in the postmortem meeting or simply shut up and move on because reality is that we are hired to do some stupid task and not to think for ourselves?
- mc32 6y agoDon’t bite or embarrass the hand that provides you make-work... Seriously they probably tested it and it worked in theory, just not in practice and now they fix it for reals.
- acruns 6y agoI guess if their DR firedrill assumed their failover router hosting the status page could never go down it would pass, but come on IBM.
- detaro 6y agoSeems a common mistake. If I remember right, AWS stumbled over their status page depending on S3 a few years back
- mbreese 6y agoI view this as growing pains that everyone has to learn the hard way. The bigger question is will they learn the lesson and how will they fix this for the next time? (Because there is always a next time)
- fogetti 6y agoWhile that is a solid advice, still it raises the question: is there such organization which corrects itself by firing the incompetent and promoting the competent instead (the "toldyouso" guy in this case)? Or we just simply accept and making it the norm that even the lowest level of organizational governance is corrupt? I am serious about this, because how people perceive their own rights, their own roles, their own status, their own influence and their organization's wrongdoing will influence the attitude in the long run against each and every organization in society in my opinion. I know that I was blowing the question out of proportion, but it bugged me to ask anyway.
- stevehawk 6y agoguess they didn't learn from AWS and hosting their status pages (in particular their icons) in S3
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]
- whyleyc 6y agoIs there talk of an ETA for any fix yet?
- fuckyah 6y agoNever
- deleted 6y ago[deleted]
- rabee3 6y agoThank you
- toast0 6y agoI managed to capture a traceroute on the lg.softlayer.com from seattle to my home near seattle, that went via London (networklayer to london, then telia back to seattle). lol Looks like things are getting better now though; looking glass says seattle and dallas can get to me without going to crazy destinations. wdc is still icky.
- bigiain 6y agoDoes it smell of fuckup, or of intentional BGP hijack?
- nixgeek 6y agoIf it was a BGP peer who normally sends you 3 prefixes with under a /20 in aggregate and they suddenly started sending you a whole table, or if you allowed a peer to send you a default route, then both of those are highly avoidable through session configuration and filtering. If the route which caused the madness came in via a large settlement-free peer (like a big eyeball/access network) or a transit (which is probably giving you the whole table) that's entirely another story.
- Fordec 6y agoI remember I was at an IBM sponsored hackathon around 2015 where it was a requirement to use Bluemix. Over the course of the weekend the service went down for hours 3 times. Literally this morning I was wondering what ever happened to it, like did it die a quiet death? Oh it rebranded to IBM cloud in 2017. Now this news. I think there's an eponymous law named for this sort of thing.
- deleted 6y ago[deleted]
- kinghuang 6y agoThat's funny. I've had the exact same experience with Bluemix at a Hackathon in the past. It was down for almost the entire weekend, screwing all the teams that didn't pivot early enough.
- ComputerGuru 6y agoSo what are HNers using IBM Cloud for and where do you see that it has an edge over AWS offerings (where an overlap exists, obviously)? (I figure either you’re in devops and you are putting out fires too busy to read this thread or you’re not and your work is halted because of the incident so you might have time to read and reply ;)
- spydum 6y agoI suspect nobody really uses it outside of weird outsourced financial modelling/planning tools like TM1 and other apps people stopped wanting to manage themselves.
- freehunter 6y agoI work as a consultant with big enterprise companies and I can assure you big enterprise companies are using IBM very heavily. As well as Oracle and HP and other uncool tech companies.
- manquer 6y agoRarely it is a just technical decision, usually money is the reason. In small and mid size organizations the CSP gave better pricing, or they help with your sales etc In large organizations - IBM/Oracle bundle their existing products currently being paid for any way, or account managers have great relationships with decision makers , the company already has signed up big multi year deals. This is not just IBM, it applies to GCP/Azure/AWS as well.
- Xenoamorphous 6y agoIf we speak specifically about IBM Cloud vs AWS, we use the Natural Language Understanding API in IBM Cloud and as far as I know the equivalent AWS offering, Comprehend, doesn't provide named entity disambiguation nor links to knowledge graphs (IBM links to DBpedia). MS and Google do provide those features though.
- dsmcr 6y agoUnfortunately, that API changes regularly and often in undocumented ways that causes breakages for customers. Its really a lot of fun to deal with when suddenly a bunch of automation breaks and it turns out an unannounced push fundamentally re-writes foundational API calls.
- homeglue 6y agoI've seen multiple services get affected this morning including Sendgrid, Nexmo and Up bank, all at the same time. Wondering if this is related.
- Operyl 6y agoYup .. hit us pretty badly. Our account manager doesn't know either.
- shaabanban 6y agowonder if we'll ever get a post-mortem about this... Seems to be global
- Operyl 6y agoMaybe. About 3/4 of all outages get a post mortem. There's 1/4 of the time they refuse to tell us anything.
- colinbartlett 6y agoDo you actually have data on that or are you conjecturing? Because I would really love to see data about that if it exists somewhere.
- Operyl 6y agoI'm talking from experience. Most things do get post mortems, but there's a lot of crap they also don't give us post mortems for "because customer data." It's my number 1 complaint, and I fight with managers about this all the time. We have a ton of hypervisor problems, and a lot of networking issues (generally over private network) and they tend to get very very secretive about it.
- mbreese 6y agoIBM cloud specifically or just in general?
- Operyl 6y agoI'm talking about IBM Cloud specifically, yes.
- toast0 6y agoI didn't use their hypervisors, but I've had a lot of experience troubleshooting their networks. They've gotten a lot better at proactive monitoring, but we used to occassionally find some private networking paths that were having trouble, and until we narrowed it down, it was hard to find. (I dunno, I guess you can't just ask all the routers if there are any ports with errors, but sure enough, when they found the right port, there was usually a huge error count, or something) The key thing is each IP 5-tuple (peerA, peerB, protocol, portA, portB) will always take the same path over their network (most likely a different path for return packets, when A and B are switched), so in order to properly probe, you need to probe on a lot of of port combos, and once you find a broken combo, you need to run MTR on those ports, so you can give them the MTR that shows the issue. Or, if you can, have your internode protocol run on multiple connections and drop connections that are showing issues, and let a different customer file the tickets :) (email is in my profile if you want to discuss)
- shaabanban 6y agoAlso still no communication from IBM that anything is wrong.
- Operyl 6y agoAccount managers are texting, but they have no VPN access right now.
- mark-r 6y agoLet me guess, their VPN authorization runs on the IBM cloud.
- _jal 6y agoDogfooding is great, but you need to think it through...
- adrr 6y agoIt seems all their external network connections are down. I assume people will have to drive to the data centers to fix. I really want to see a post mortem on this outage.
- wmf 6y agoThe data centers are staffed 24/7 and out of band is also a thing.
- whyleym 6y agoFound this from a user on Twitter - "Our status page for IBM Aspera is on StatusPage, so you can track here as a bank shot: https://status.aspera.io https://status.aspera.io "
- sky_rw 6y agoThe most infuriating thing about this is the ZERO communication coming out of IBM Cloud. No emails. No updates to twitter. Status page down. Support lines clogged. At least give me something I can point my customers at to show them this is not due to my incompetence.
- bizt 6y agoYep, super annoying I had to link my customers to a techcrunch page :(
- gatvol 6y agoWell if they cannot foresee this eventuality, what else are they missing under the hood?
- nadavami 6y agoIt seems like the status page just came back up.
- blazefox69 6y agoFixed it for you https://github.com/ibm-cloud-docs/overview/pull/74 https://github.com/ibm-cloud-docs/overview/pull/74
- ck2 6y agoeven weather.com was down but someone broke ebay too Fastly error: unknown domain: www.ebay.com. Please check that this domain has been added to a service.
- toast0 6y agoweather.com makes sense. IBM bought the weather channel a while ago, hosting is likely tied to IBM Cloud at this point (although it looks like it's fronted by Akamai)
- vmh1928 6y agoIBM bought the technology part called the Weather Company. That's the part that gathers weather info from all over and makes it available. The cable TV channel is still independent.
- supernova87a 6y agoAha, I guess explains why Wunderground.com was out too.
- pmarreck 6y agoImagine hosting your status page on a different domain
- 9nGQluzmnq3M 6y agoDNS worked fine here, this was an infra issue.
- deleted 6y ago[deleted]
- someguy12321 6y agoheads be rolling tomorrow!
- deleted 6y ago[deleted]
- kitteh 6y agoAbout a month ago their Northern Virginia region was down. All the BGP prefixes associated with it disappeared from the internet (routes withdrawn). This time (I went to check when someone mentioned it) they kept advertising, but all traffic went nowhere once it got into their network. Curious to see if there is an RFO released.
- aiisjustanif 6y agoI wish we had a record of this.
- kitteh 6y agoI do. I store all this stuff. Where should I put it?
- bantec 6y agoIt’s a second significant issue for last year with IBM( absolutely inconsistent for critical infrastructure (we are FinTech)
- anon102010 6y agoA quick check of cloudflare's isbgpsafeyet page IBM Cloud - unsafe At least AWS signs their routes I think. If you can't even sign your own routes - hard to have a ton of pity.
- kortilla 6y agoSigning routes doesn’t mean others reject unsigned routes. AWS is just as vulnerable to hijacking as anyone.
- vmh1928 6y agoIn the Cloud Status History page scroll down to the 6:32 entry that says "Unable to Access IBM Cloud" https://cloud.ibm.com/status?selected=history https://cloud.ibm.com/status?selected=history - 2020-06-10 02:19 UTC - RESOLVED - The network operations team adjusted routing policies to fix an issue introduced by a 3rd party provider and this resolved the incident
- akerro 6y agoHaha, amazon had the same problem a few years ago when they had fire in datacenter, their status checker page was hosted in the same building and was showing everything is fine, while 1000s of websites hosted on AWS were down.
- nonines 6y agoThis looks related (smoking gun?) https://status.aspera.io/incidents/t9r03x71dxkl https://status.aspera.io/incidents/t9r03x71dxkl >> A 3rd party network provider was advertising routes which resulted in our WW traffic becoming severely impeded.
- rbanffy 6y agoIt can only be attributable to human error. No IBM computer has ever made a mistake or distorted information. They are all, by any practical definition of the words, foolproof and incapable of error.
- deleted 6y ago[deleted]
- oehpr 6y agoThanks for keeping us up to date. The timing on that incident was a weird coincidence for our team. We had just rolled out a bunch of production updates, and then 5 minutes later ALL our VM's went down and I freaked out and tried to diagnose what was happening and recover before I started noticing all the other people going down. So. Hate to ask this. But what happened to your old posts that are now just "-"'s? I mean I can guess. It's just that the communication blackout while the incident was happening was not appreciated. And you stepping forward to let us know what was going on was very appreciated.
- deleted 6y ago[deleted]