7 ms·
To others wondering what this company does: > Ably is a Pub/Sub messaging platform that companies can use to develop realtime features in their products. At t
by smashed 5y ago
To others wondering what this company does:
> Ably is a Pub/Sub messaging platform that companies can use to develop realtime features in their products.
At this time, https://status.ably.com/ https://status.ably.com/ is reporting all green.
Althought their entire website is returning 500 errors, including the blog.
It is very hard not to point the irony of the situation.
In general I would not be so critic, but this is a company claiming to run highly available, mission critical distributed computing systems. Yet, they publish a popular blog article and it brings down their enire web presence?
- konschubert 5y agoThey're not selling a blogging platform, as long as pub/sub works it's fair to say that they're up.
- rvz 5y agoWell said. Their main service and offering is still up. They are not hosting blogs or websites as their business. This reaction is a storm in a tea cup left inside of a sunny cottage. Downvoters: I don't see anyone screaming at this stock trading app that also stopped using Kubernetes [0]. [0] https://freetrade.io/blog/killing-kubernetes https://freetrade.io/blog/killing-kubernetes
- mountainriver 5y agoThis is the face of your company, if you can’t handle the increased load on your blog why should I trust your other systems. I run on k8s and it works really well
- Nicksil 5y ago>This is the face of your company, if you can’t handle the increased load on your blog why should I trust your other systems. I suspect the sort of individual visiting this website and reading the blog will know that their services and blog do not run on the same machines.
- mountainriver 5y agoI suspect that when someone’s website is down it looks bad on them
- rvz 5y agoI would have agreed with you if 'their blog AND their main service' also went down; Not just the website. But their main service is still up. 95% of the comments here parade about a hiccup of this service having a blog going down when that is not even their main service but continues to use an unreliable service like GitHub even when that goes down every month and they also don't use K8s either. GitHub (and GitHub Actions) was proven to be very unreliable and the whole service went down multiple times and somehow that gets a pass? Even GitHub's status page returned 500s due to panicking developers DOSing the site. Same with GitHub pages. But a single blog getting 500s but the whole service didn't go down? Nope.
- gbe52 5y agoIt’s to be expected. Criticism of Kubernetes is taken personally by people who have dedicated themselves to it since it’s in vogue. It is very much showing the cracks in its perceived perfection and this company is far from the only one bucking the trend. We are seeing the long tail of people who don’t realize (or acknowledge) that Kubernetes is already on the other side of the hype cycle. That long tail is currently driving this conversation because commenting on a 500 is the lowest effort imaginable, but they do not speak for the industry. (Source: 24 years in FAANG-scale ops. I remember this cycle for many, many other fetish tools.) Seriously, a Wordpress blog goes down (like that’s never happened under HN load) and all of HN is saying “lol what dogshit if only they used k8s for basically static content, the morons” which tells you a lot about the psychology going on here. A LOT. Otherwise smart engineers just can’t help themselves and redefine “irony” even though it contributes absolutely nothing to the discussion and doesn’t even address the fundamental criticism. We are in a thread that essentially started with “the blog is down, how can I trust the product?” If you can figure out that logic please let me know. That’s basically saying “I don’t understand operations at all,” but here we are, listening to it and its Strong Opinions on a resource allocator and orchestrator. This thread basically confirmed for me that Kubernetes is losing air. I already knew that, but the signal is getting a little bit stronger every month.
- foobiekr 5y agoVery well said and agreed.
- foobiekr 5y agoAuto scaling to cover insane unpredicted load like this is t really representative of anything other than a failure to have cost management in place.
- foresto 5y ago> This is the face of your company, if you can’t handle the increased load on your blog why should I trust your other systems. https://xkcd.com/932/ https://xkcd.com/932/
- remram 5y agoThey are selling a highly-available product. You don't pay for pub/sub, anybody can run that, you pay for reliability and scalability. Their approach to other parts of their infrastructure definitely can and should inform your decision to buy into their product. In addition to insight in their infrastructure practices, this gives you a unique opportunity to look into how they deal with outages, whether they update their status page honestly (and automatically), and how fast they can solve technical issues.
- _jal 5y agoIt is embarrassing. It also highlights a growing pet peeve of mine: the uselessness of status dashboards, if not done very well. They're much harder than they look for complex systems. Most companies want to put a human in the loop for verification/sanity check purposes, and then you get this result. Automate updates in too-simplistic a way, and you end up with customers with robots reacting to phantoms.
- jrockway 5y agoStatus pages are designed to be green so that Sales can tell people you're never down. Everyone does it, so you have to do it too. My experience with monitoring 100% of outgoing API calls is that many services have well below a 100% success rate. Sometimes there are just unexplained periods of less than 100% success, with no update to the status page, and sometimes there's even a total outage. (I had a sales call with one of these providers, and they tried to sell me a $40,000 a year plan. I just wanted their engineers to have my dashboards so they can see how broken their shit is.) The one shining light in the software as a service community is Let's Encrypt. At the first sign of problems, their status page has useful information about the outage.
- intev 5y agoThe "Website portal and dashboards" is picking up the disruption now.
- lallysingh 5y ago.. and Kubernetes does health monitoring!
- mountainriver 5y ago….and autoscaling
- deleted 5y ago[deleted]
- leephillips 5y agoThe HN post could have brought down their website, but do you know that for a fact? Maybe it was some unrelated incident or attack.
- mLuby 5y agoI see their status page showing red for the website. https://imgur.com/a/mBbfLtZ https://imgur.com/a/mBbfLtZ I'm not surprised when a company's status page only reports on (highly available) services, not the website or blog—which are likely run by marketing, not engineering. Still, it's simple and free to set up status pages so there isn't much excuse.
- caeril 5y ago> Althought their entire website is returning 500 errors, including the blog I've been here in a similar situation, and my guess is they've: 1. Reverse-proxied /blog to a crappy Wordpress instance run without caching. 2. HN traffic killed the blog. 3. Cloudflare, in their infinite wisdom, if they see enough 5xx errors from /blog, will start returning 5xx errors for every uri on the site, with some opaque retry interval. 4. Voila, the entire site appears to be dead.
- judge2020 5y agoCF doesn't do 3) to my knowledge.
- PaywallBuster 5y agoHighly doubt it they'd do it
- btown 5y agoMore likely the entire www. server is a CMS that fell over, and the application is on a separate subdomain, and they are only monitoring the application.
- babelfish 5y agoCan you detail (3)? Is this a product feature documented anywhere? As others have pointed out, this does not seem likely
- throwaway_2047 5y ago99.99% SLA achieved. Jokes aside, I think the article is well written.
- paddybyers 5y agoAbly CTO here. Well that went well ... 1) The realtime service and website are different things. The blog post is talking about the service, which has been continuously available. 2) Oops, the website fell over. We'll fix that. Thanks for all the advice :)
- m0nst 5y agoAs a ton of engineers have transitioned to kube already. Like years ago, you have just made ramping up new engineers to your custom setup a pain point for scaling. But you probably already knew that. =)
- youngtaff 5y agoIf you're running the public website through Cloudflare, have a look at APO makes caching WP at the edge easy
- nrmitchi 5y agoIf a marketing website/blog is not the thing you are selling, it very often does not make sense to host it yourself. Your time and energy is better spent focussing on other things, and paying a separate hosting/marketing company to manage that. I can't say for certain that this is what happened here, and the irony is definitely there, but overall it's not valid to judge a company's reliability on a piece that is not what they are selling, and are likely to not even be managing themselves.
- void_mint 5y agoIs your assertion that if they were on Kubernetes, this wouldn't have happened? What's ironic about a blog going down from a company selling pub/sub software?
- handrous 5y agoBesides, though the situation's undeniably funny, there are much more appropriate, simple, and proportionate ways to ensure your WordPress marketing site or blog don't fall over than putting them on k8s, even if you are a k8s shop.
- ren_engineer 5y agothis implies their actual service and marketing website are running on the same infrastructure, which is unlikely
- NicoJuicy 5y agoIt's stated on the status page: > Our blog and website are experiencing load-related issues, leading to slow loading and 5xx errors. All backend services (realtime and REST apis) are on entirely separate infrastructure and are unaffected.
- jesterson 5y ago> In general I would not be so critic, but this is a company claiming to run highly available, mission critical distributed computing systems. Yet, they publish a popular blog article and it brings down their enire web presence? You may expect that company blog have a bit lower uptime than services it offers. As someone on purchasing side, I don't give a f** if a company website or blog is down for weeks, if their services are operational. As a disclaimer - never heard of Ably, but wholeheartedly support non-kubernetes environment. Being CTO in large e-commerce company we do not use or even plan to use kubernetes.