5 ms·
What use are status pages when you need to get VP+ clearance to acknowledge an outage? We see this time and time again with major services. It is a shame that
by juice_bus 4y ago
What use are status pages when you need to get VP+ clearance to acknowledge an outage? We see this time and time again with major services.
It is a shame that we have to find out about outages on HN.
- toomuchtodo 4y ago> What use are status pages Marketing and sales. Not /s We are in dire need of something crowdsourced, or where someone like DataDog or other telemetry systems offer you the ability to share non sensitive metrics publicly for various cloud or SaaS systems that they publish. Edit: y’all are amazing with these monitoring tools!
- austinpena 4y agoBuilt something like that at https://taloflow.ai/is-aws-down https://taloflow.ai/is-aws-down Looks like there’s been some errors too
- CoastalCoder 4y agoDoes downdetector.com meet that need? At least for Steam, I've found it pretty useful.
- toomuchtodo 4y agoFor public facing endpoints or front ends, yes. For complex systems where you’d need sensors inside (AWS, GCP, anything IaaS or PaaS, etc), no. Somewhat but not entirely similar to BGP looking glass systems.
- geocrasher 4y agoYou mean like https://mcbroken.com/ https://mcbroken.com/ ?
- WarOnPrivacy 4y agoOooooo. OpenStreetMap has clearly visible county lines. I'm in lust.
- i67vw3 4y agoDowndetector was itself down today for half an hour when the cloudflare incident happened.
- mbesto 4y ago> We are in dire need of something crowdsourced, or where someone like DataDog or other telemetry systems offer you the ability to share non sensitive metrics publicly for various cloud or SaaS systems that they publish. This is literally what I'm building right now. See reply above: https://news.ycombinator.com/item?id=31825239 https://news.ycombinator.com/item?id=31825239 Shoot me an email if anyone is interested in getting beta access.
- jmartens 4y agoThat's exactly what we are building at https://metrist.io/ https://metrist.io/
- cmccart 4y agoNew Relic has had this for a few years now for a couple hundred of the most requested domains: https://docs.newrelic.com/docs/query-your-data/explore-query-data/dashboards/explore-public-api-performance-dashboard/ https://docs.newrelic.com/docs/query-your-data/explore-query... Disclaimer - I work at New Relic but not on this.
- jabroni_salad 4y agoI use uptimerobot to monitor a lot of endpoints that I depend upon but don't really control. Been burned by first party status pages way too many times.
- lamontcg 4y ago> Marketing and sales. Not /s When I got hired at Amazon in 2001 we had a "gonefishin" page that was a static page that would be served in the event of an outage (this was before status pages, but it was kind of the same thing -- public acknowledgement of a major incident). The standard protocol was within minutes of a sev 1 to make a decision to display the GF page once it was confirmed that the whole site was down and then work to fix the issue. By the time I left in 2006 that was no longer policy since reporters had setup monitoring for that page to detect outages and report on service availability so they just let it crash and return 500s or whatever the failure mode was. Optimize for making the job of external agencies doing reporting on their availability harder instead of easier.
- deleted 4y ago[deleted]
- philote 4y agoI guess one could argue it isn't an outage since it only seems to have affected a subset of users. I got on a zoom call when this issue started and we had 3 of the 4 participants. Only one couldn't connect due to the issue. But I do agree they should be able to monitor things better and show some sort of update on their status page as soon as possible.
- Wowfunhappy 4y agoI feel like this is almost worse. It would be awful if you were the only person who couldn't connect to a high-stakes meeting. At least if it happens to everyone, it's obvious that the problem is on Zoom's end.
- danachow 4y agoIf you stake your life’s happiness on pleasing morons (and the morons in this case are those that pretty much don’t immediately assume technical problems out of your control) - you’re pretty much guaranteed a bad time.
- Wowfunhappy 4y agoA couple of months ago, I finally landed a first-round job interview at a place where I've wanted to work for several years. The interview was conducted over Zoom. What would have happened if Zoom had worked fine on their end, but I was randomly unable to connect? Perhaps it would have been fine—they would have been understanding, and we would have rescheduled for another day. Perhaps if they hadn't been understanding, I shouldn't have wanted to work for them anyway. But, I don't know. I wanted to work for them, and I was competing with other candidates who presumably interviewed on different days. Hiring processes are inherently imperfect, and lots of things can be consciously or unconsciously treated as a red flag. (And yes, lots of other things could have happened on the day of the interview. But I still find this scenario particularly scary to think about.)
- danachow 4y ago
- mbesto 4y agoI'm actually building a product to solve this. If anyone is interested in beta testing, we should be rolling this out in 2~3 weeks. Shoot me an email: mbesto @ gmail service
- no_wizard 4y agoThis has made me realize why companies like pingdom have a business. I've always wondered, in the sense that I couldn't quite understand why you'd pay for someone just to ping things and alert you of outages (this was early in my career) But over the last 4 years specifically I not only understand it I can't imagine not having a service like it. Disclaimer: I don't work for pingdom and my current company doesn't use their services, I have in the past, they're pretty good, but I'm just using them as an example here
- jmartens 4y agoIf you like pingdom, you'll love what we are building at Metrist https://metrist.io/ https://metrist.io/