6 ms·
I like how quite a number of peoples answers to the on-call programmer blog was "you need better tests" here's a what if scenario:- - you have a third party s
by malbs 14y ago
I like how quite a number of peoples answers to the on-call programmer blog was "you need better tests"
here's a what if scenario:-
- you have a third party service your systems rely on
- at 4am on Sunday morning said 3rd party service upgrades their system, introducing a breaking change, having never bothered to notify users
- you get a call as the on-call person saying "application X is not longer working, please resolve"
How do tests stop that scenario from happening? Tests don't magically help you invent features/work around introduced issues in 3rd party systems.
Those are typically the on-call issues we deal with (we're on a weekly rotation)
- Smudge 14y agoI saw only two comments regarding "more" or "better" tests. One that seemed like a sarcastic low-blow ("I’m glad to see 37signals post from last week about their minimal approach to testing is keeping the new hires busy and up late at night tracking down bugs"), and the other being "Alice" who quickly got labeled a troll. Seems like both you and DHH are putting words in people's mouths.
- malbs 14y agoI'll wear that, but also I just happen to be on support this week so it's a raw subject right now ;)
- anthonyb 14y ago> Tests don't magically help you invent features/work around introduced issues in 3rd party systems. Uh, yes they do. You want a unit or system test which covers the case where an external system is down or returns something that you can't parse. Something like: # code to take third party thing down # eg. mock out lib and return nonsense (unit tests) # or add an /etc/hosts entry (system tests) assert "Sorry, but that feature is unavailable." in page.content Now the entire app doesn't asplode, and you can wait until 9am to fix it. Follow up is to make sure that you're on whatever mailing list tells you when changes are coming. The only case that this doesn't cover is when it's a) an essential part of your app, which b) you aren't paying for and c) they don't have a mailing list, in which case wtf? you need to find a better 3rd party library/service. ps. Look up the "chaos monkey" - it's very enlightening :)
- jedberg 14y ago(I work for Netflix) It's funny that you mention the Chaos Monkey, considering that Netflix has 24/7 on call programmers for tier 1 support. We do however also make great efforts to make sure that we are resilient as possible to failure of 3rd party services.
- anthonyb 14y agoI suspect you also pay for your 3rd party services, which the GP's company doesn't seem to do.
- malbs 14y agoa, b, and c, and there is no other / better service, so it's a difficult one to solve. also things can't wait until 9am or the selected waking hours, we have too many people/systems relying on working infrastructure, so an issue popping up at an ungodly hour is fixed there and then, even if it means calling other people. Those are the worst support calls, 3am on some weeknight, and you can't actually fix the problem because you're not 100% certain, and you need to call a colleague and wake them up too. You feel like an asshole.
- anthonyb 14y ago> a, b, and c, and there is no other / better service, so it's a difficult one to solve. You don't say what the service is, but if your company is relying on it to the extent that you need to be awake at 3am, then I suspect it's well worth calling the people providing it and offering to throw money at them. Otherwise you're essentially relying on their goodwill for business continuity...
- v21 14y agoI used to work support for a SMS aggregator. The core business was delivering texts to mobile networks. Mobile networks break all the time, it turns out. And can't really be replaced - there's only one T-Mobile. Now, when they died or sent garbage, our software wouldn't crash. But our service would stop functioning. And alarms would go off, and we would have to confirm why, and call them, and ask them if they knew if it had stopped working and why. (The follow up was indeed to always ask to be added to the mailing list that would let us know this ahead of time. These proved remarkably unhelpful. In once case, we ended up setting up having to set up a mobile number to receive SMS alerts - this was apparently the only way they would notify anyone?) And reasonably often actual engineers would be needed to be woken up. And in several cases, change the parsing behaviour so we could handle sudden unexpected changes in the format they returned. Yes, in the middle of the night. Our system wasn't perfect, but I'm not sure our problem was simply a lack of system or unit tests. During the day, we largely dealt with more minor customer complaints, and ongoing maintenance, outages and other nonsense fromt he carriers. Engineers would often have to dig into these edge cases too, and there was a nominated maintainer to look at this stuff, to let the rest of the team add features and work on long term fixes. But sometimes you just need someone to wedge LargeCustomer's encoding settings because they can't figure out how to properly specify it on their end.