6 ms·
It seems like most work done to make distributed systems reliable is aimed at handling machines or groups of machines going down (e.g. the leader node in one re
by piinbinary 10y ago
It seems like most work done to make distributed systems reliable is aimed at handling machines or groups of machines going down (e.g. the leader node in one region goes down at the same time as an entire other region). This half seems to be a solved problem.
The postmortems published by Google, Amazon, and Azure (as well as postmortems internal to the company I work for) are nearly always due to some type of change (code or configuration) being rolled out. It seems to me that we need some help from computers to make these systems reliable. - something like static type checking in a programming language, but applied to a distributed system. Perhaps your architecture wouldn't "compile" if the network traffic will go the wrong place, or if a rate limit is above the capacity something is expected to handle, or if the change would impact too many servers at once.
- foolfoolz 10y agoat the platform-wide incident level its more often configuration than code
- BinaryIdiot 10y ago> Perhaps your architecture wouldn't "compile" if the network traffic will go the wrong place, or if a rate limit is above the capacity something is expected to handle, or if the change would impact too many servers at once. This is an interesting idea. I'm always terrified when I have to deploy a minor configuration change into a production system that gets distributed out; if there was an easy way to apply some sort of check against all of it (you know without having to build something explicitly to do this) that would be awesome. I wonder if the future of something like this might be using containers. Each piece that gets deployed is in its own container networked together and you stand it up in a staging environment, ensure it's talking together correctly through some sort of set of integration tests then push it into the production environment.
- lima 10y agoIn many cases you can do staged or partial rollouts, observe error rates or whatever suitable metrics you have, and then continue with the next bunch of servers.
- DigitalJack 10y agoI wish the law could be statically checked. In the US at least you are, at any given moment, likely breaking a law because for each and every law, there is one that contradicts it (at least partially).
- tener 10y agoQuite likely you'd first have to invent strong AI which would give you a chance of defining a coherent theory of law which is really a tangled mess. Pretty sure no human has capacity to do it on their own in the foreseeable future.
- DigitalJack 10y agoI think I'd approach the problem from the other end, find a way to write laws such that they are checkable with assertions, properties, and constraints.
- BinaryIdiot 10y agoThis would be amazing. If we had this the first line of judges could just be automated. Feed evidence in, computer finds you guilty or innocent, done. Then if you appeal you can seek a human judge in case judgement / exceptions need to be made (so say you technically broke a law but it was actually necessary / a good thing then you have recourse). First appeals could be very quick, less traffic to first human judge, etc. Hell even without automated the justice system just doing as you suggested alone would be amazing. I've often wondered about putting laws in GIT or similar.
- plttn 10y ago> Feed evidence in, computer finds you guilty or innocent, done. Would you be willing to trust a program with your sentence? I wouldn't in the slightest.
- BinaryIdiot 10y ago
- deleted 10y ago[deleted]
- ben_jones 10y agoI'd be happy if their was just a decent linter for all configuration files. My test process right now is to spin up a vm and test something like an nginx reload or an ansible-playbook, then if it passes apply it to a distributed testing environment and then if it passes consider applying it to production. This seems crazy to me when intelli-sense for many languages can seemingly read my mind but configuration is in many ways still a trial and error process.. Does anyone know of ways to parse many of the common configuration file formats? From linting to code completion to best practice (or worst practice) checks. IF not and somebody wants my money yesterday, theirs your startup.
- jacobbudin 10y agoIf you have nginx installed, you can test nginx configurations like so: $ nginx -t -c <configuration file>
- icebraining 10y agoAnsible also has --syntax-check and even --check, which "tries to predict some of the changes that may occur".
- ben_jones 10y agoI still have to spin up the vm though to make sure changes to things like proxy settings won't affect me further down the stack. But thanks for the tip!
- notyourwork 10y agoGood idea in theory but not sure how it pans out in practice. Defining the wrong place would require some configuration would could as easily be mis-configured. Abstracting the problem to another layer isn't always the answer.
- sumitgt 10y agoI work on a service with high OLA requirements. I'm terrified everytime I roll out a new build or config change, in spite of having tested them thoroughly. At the scale at which these things operate, it's very difficult to manually imagine all the possible ways things can go wrong. Especially when you have multiple services that are all interdependent on each other and written/maintained by different teams in different geographic locations. I'm sure there is lots of research happening in this area within Microsoft and Google. Can't wait to have this sooner.
- superuser2 10y ago>The postmortems published by Google, Amazon, and Azure (as well as postmortems internal to the company I work for) are nearly always due to some type of change (code or configuration) being rolled out When you have a really large distributed system that's primarily running on metal, smaller-scale copies of the whole system per developer or even a single companywide staging environment that mirrors production are really hard, and they don't always exist. Developers work on their components in isolation by mocking out the rest of the system, hopefully there's a thorough code review, and then "integration testing" happens by flipping a feature flag and watching the logs/metrics in production. You might design the feature flagging so it initially only hits test accounts, but some things (like service communication layers, Puppet configs, router configs) don't work that way. The cost of outages resulting from the lack of a staging environment may well be less than re-architecting production (100% automation, no snowflakes, and probably some kind of IaaS) to allow for disposable dev/test environments which would exhibit the same bugs.
- adrianratnapala 10y agoBut I think what you are saying actually is an argument for @piinbinary's idea of "static analisys". If you can't really test your change, then the other thing you can do is try to rationally analyse its consequences. After all thats why "hopefully there's a throrough code review". Wouldn't it be nice to automate some of that analysis? I have no idea how to go about such a thing, or whether it is even possible -- but the harder testing gets, the more attractive this option is.
- whatupmd 10y agoThey switched to openflow on internal networks for this reason, deterministic: https://www.youtube.com/watch?v=FaAZAII2x0w https://www.youtube.com/watch?v=FaAZAII2x0w External network is BGP though and sounds like they didn't detect it until user reports hit. They can't predict the problem and there detection isn't working well either.
- asuffield 10y ago(Tedious disclaimer: my opinion only, not speaking for anybody else. I'm an SRE at Google. My team is oncall for this service and I know exactly what happened here; I probably can't answer most questions you might have.) > Perhaps your architecture wouldn't "compile" if the network traffic will go the wrong place, or if a rate limit is above the capacity something is expected to handle, or if the change would impact too many servers at once. So in the first instance, I tend to like this sort of idea. However: we are already substantially ahead of the sort of things that you're thinking of. Full static simulation of a system as complicated as all the components involved here is... well, I can sort of see how it could be done, but it would be a herculean effort; I don't think it would ever be good enough to catch cases like this the first time they happen. There are systems where this sort of thing can be done, but all the ones I can think of are much smaller in scope.
- rixed 10y agoWhat does "static simulation" mean ? Probably not static analysis.
- asuffield 10y agoTo fully answer questions like "how much traffic will go in this direction?" you need your analysis to include a simulation of what the entire internet is doing. That's hard. I can't talk about the details, but you can assume that "static analysis" of the form being talked about here is something we've already done, and it's not enough to handle cases this complicated.