4 ms·
We've taken steps to minimize outages as much as possible. The system is distributed across 3 data centers, with fast automatic rollover in case of a data cente
by alexsolo 16y ago
We've taken steps to minimize outages as much as possible. The system is distributed across 3 data centers, with fast automatic rollover in case of a data center outage. We've architected the system to ensure we never drop alerts. PagerDuty integrates with monitoring via email or API; if we receive the message on our end, we guarantee you will be alerted. We've had a few incidents where we have delayed sending out the phone call or SMS alert for a few minutes, but we've never dropped an alert.
In terms of setting a formal SLA, we haven't done so mainly because we're not sure how to go about implementing this. I've checked the SLAs of a few hosting and cloud providers including AWS, Rackspace, Linode and Slicehost, and I haven't found a compelling example to work from. Some of these guys don't have an SLA (they try their best) and the others give you only a portion of your money back.
The whole point of an SLA is to incentivize us to never go down. In our case, we know that if we ever go down, we will lose our customers; that's incentive enough :). Having said that, we may still add an SLA guarantee as part of a larger "enterprise" pricing plan.
We definitely plan on adding plugins for all the popular monitoring systems. We've also released an integration API to allow PagerDuty to integrate with any system that can make an HTTP API call (or call a command-line script that can do this).
I'm pretty sure Zabbix will work with PagerDuty right now, via the integration API. We'd love to work with you to set this up. Please send me an email at alex@pagerduty.com.
- alexsolo 16y agoI'd love to hear what some of you think about SLAs. Is it worth implementing one?
- mahmud 16y agoAt a former employer, it took weeks to negotiate the terms of an SLA with a solution provider and it ultimately fell through. We glossed over several vendors because of the lack of one (it's usually negotiated, not pre-written like a privacy policy or T&Cs.) However, what ultimately made the deal was not an SLA, but a new vendor that showed substantially deeper pockets than we expected. A meeting was organized with their sales guy and FOUR suits showed up in a black sedan. Most gangsta display of power, and our glass cubicles were gassed down with the musk of cigar, Brut and Drakkar Noirs.
- btilly 16y agoThere are lies, damned lies, and SLAs. Personally I only find an SLA useful if it is worthwhile. Most of the SLAs out there aren't. And for good reason. You should probably offer one, but like a smart company shouldn't make the burden too bad. Suppose someone doesn't respond to a page. Is it because they were too far asleep to hear the paging device? Because the paging device didn't work? Because some other problem kept them from working on the page remotely? Because their carrier blocked the page? Because you broke down? Because the problems in their system kept them from sending you the information in the first place? There are a lot of points of failure. And your service is not one of the more likely ones to break. Furthermore if there is a dispute, whose records win? They didn't respond to a page, your records say they never sent the page. They blame you, how do you resolve that? Therefore I'd suggest offering an SLA, but make it be something like, "If you missed a page and are convinced that it was our fault, we'll refund the last X months." From your point of view it is a no questions asked refund policy, that carries with it the consequence that that person is not allowed to sign up for your service. (Unless, of course, you're convinced it was your fault they didn't receive their page.) But whatever you do, be careful not to accept potential liability for something that likely was their problem. I would also suggest that you share best practices. For instance an important one is that companies need to provide a well-defined escalation path. Recognize that humans fail (whether because of not waking up, being in the process of driving, etc) and so people are unreliable components that need a fall-back mechanism. The act of educating your clients about things like this will help them avoid problems that could cause them in an imperfect world (ie the one we live in) to become unhappy with you.
- lsc 16y agoSLAs with exceptions based on "fault" are meaningless. Either you guarantee you will keep your shit working, or you don't. (Either way is fine, really... but arguing over "fault" is not a productive activity.)
- mseebach 16y agoIt's not meaningless? If "working" is dependent on several pieces working, and only some of them is under your control, you can be in a state of "not working" without being at fault. I've had a server go down for a large group of users because of a malconfigured routing table between them and the server. If we'd had an expensive SLA, there would have been significant "what the heck is it we're paying for, then?" discontent.