2 ms·
That's a funny anecdote and a good step, but you guys still have a long way to go in regards to outage alerting at SendGrid. I was annoyed when after a several
by BluePen7 6y ago
That's a funny anecdote and a good step, but you guys still have a long way to go in regards to outage alerting at SendGrid.
I was annoyed when after a several hour long outage SendGrid still had not updated their status page to reflect an issue.
But I was even more annoyed when 5+ hours into the outage, a different support person told me they aren't aware of any outage because their status page does not reflect one.
If I sound bitter it's because I am, I was the on call person who got woken up in the wee hours of the morning to deal with this. I'm referring to the outage you guys had in December if you're wondering.
- sethammons 6y agoReal sorry to hear that, and surprised. I’ll look into this. If you wouldn’t mind emailing me more details, that would help.
- zenexer 6y agoI had the exact same experience. I was on vacation with family, but was on call. There was a bug on SendGrid’s end that only affected a small percentage of emails; in particular, they had the wrong DKIM signature, if I remember correctly. I spent the entire vacation trying to convince SendGrid engineers that it wasn’t a configuration issue on our end. They were rather rude and kept insisting that it wasn’t an issue on their end. Eventually we gave up and migrated away from SendGrid. If I remember correctly, they did fix it, at least partially. I think it had something to do with one of their prod servers running a dev config, resulting in DKIM signatures for a SendGrid-owned test domain—something along those lines. That may have been a different incident, though. That was just the straw that broke the camel’s back. There were often issues that only affected one or two servers—we were often told that a server didn’t have an up-to-date configuration or something like that. Each time, it was the same process: convince them it wasn’t an issue on our end, escalate it repeatedly until it lands on an engineer who understands the issue, they fix it in an hour after wasting days of my time insisting that yes, I know what a DNS record is. And the whole time, some (small) percentage of emails are failing. I can only provide so many logs, DNS records, and email headers to demonstrate the issue; at some point, someone has to be able to look at it and go, “well, yes, that’s very obviously the wrong value, and nothing a customer sets on their end should ever have that effect.” Sure, Slack dropped the ball here, but at least they updated their status page. Even for small issues, when I contact them, they either get it resolved or tell me it’s on their to-do list—no need to convince people that I understand basic concepts.