3 ms·
I'm attempting to even imagine how one would build a useful way to test this. Would they have to have a secondary, world-wide datacenter network with all their
by mjibson 10y ago
I'm attempting to even imagine how one would build a useful way to test this. Would they have to have a secondary, world-wide datacenter network with all their various services behind it?
- ikeboy 10y agoYou could have it send messages to the actual servers, but with an added flag that says "fake", which makes the servers ignore the message/send back a message saying pass/fail/whatever (testing the flag could happen first, one server at a time manually). Then check whether the program continued to push updates.
- maxander 10y agoYou may be able to build an elaborate system of dummy network operations to test with, but this system may wind up with bugs that mask what would be errors in the real system. And how to you test against that? A dummy network to test the dummy network operations on? What if the dummy network contains bugs that make it behave significantly different from the real network, in error cases? How do you test for that? Its turtles all the way down!
- ikeboy 10y agoYou can test it against the actual network; if something goes wrong, you'll have downtime, but you'll be prepared to get it all back up. Or, to test whether the "prevent errors from going to new places" works, temporarily configure the new places to ignore new configs; if the system works, no messages will be sent there; if the system doesn't work, they ignore the message and you learn about a bug.
- dward 10y agoYou mean an ICMP request? The IPs were anycast and did not become unreachable until all edge routers had stopped announcing BGP routes. At that point the failure was global. Check out the postmortem, it's a good read.
- ikeboy 10y agoThe servers became unreachable, but the IPs weren't unreachable until all the servers were reconfigured. My test should have caught this bug: > In this event, the canary step correctly identified that the new configuration was unsafe. Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process, and thus the push system concluded that the new configuration was valid and began its progressive rollout.
- manquer 10y agoWhile testing would have been quite difficult, any simple canary release or timed release mechanism would have prevented this / limited the damage. At such mission critical systems, applying any global change in a such manner is asking for it, Devops can also be SPOF, this seems one such case.
- mgw 10y agoThey had a canary release mechanism in place. This is described in the post mortem. > These safeguards include a canary step where the configuration is deployed at a single site and that site is verified to still be working correctly, and a progressive rollout which makes changes to only a fraction of sites at a time, so that a novel failure can be caught at an early stage before it becomes widespread. In this event, the canary step correctly identified that the new configuration was unsafe. Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process, and thus the push system concluded that the new configuration was valid and began its progressive rollout. Taking no cofirmation of the canary testing process as a signal to go ahead though is not just a bug but a design flaw IMO.
- magicalist 10y agoThere was a canary release.
- mattdeboard 10y agoIf you read the actual report, it mentions that they did a canary step but its effectiveness was undermined. > In this event, the canary step correctly identified that the new configuration was unsafe. Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process, and thus the push system concluded that the new configuration was valid and began its progressive rollout.
- awinter-py 10y agobetter test -- if you have seldom-exercised edge-case functionality that you can't figure out how to test, remove those features.
- Rapzid 10y agoYes, in a manner of speaking; physical or virtual lab. At googles scale it wouldn't be unreasonable to have a completely parallel, but scaled back, network where they test their automation and code for happy and sad path. That doesn't mean that bugs can't creep in. Who knows, maybe these were all extremely unlikely bugs and Google hit an astronomically unlikely bad-luck streak. Happens.
- pmarreck 10y agoYou fake out the connection with a faker object and give that to the code that wants to communicate to the network, and it returns streamed, deterministic data that would have been expected from the actual network, given deterministic inputs. The test uses the fake; the production code gets given the real object.
- deleted 10y ago[deleted]