4 ms·
The coolest thing I took away from the SRE book was this progression of system operations from manual, to scriptable, to automated, to a fourth category I hadn'
by mikecb 10y ago
The coolest thing I took away from the SRE book was this progression of system operations from manual, to scriptable, to automated, to a fourth category I hadn't even known existed: autonomous. The idea that you can keep moving up this hierarchy of exception management beyond even chef and puppet, and systems will be able to heal themselves, is a pretty cool one.
As a manager, this made the concept of 20% time a lot more clear. These are people with the knowledge and incentive to build a hierarchy of systems that progressively remove risk from their work. This is in fact their primary business objective. And we need to make sure they have time to do that, vs working them to death with manual remediation. It's a great lesson.
Incidentally, Stackdriver contains a simple alerting and incident management tool that's really nice to use. Hopefully it gets more robust as time goes on and larger and more complex orgs move to their cloud.
Edit: not Outalator.
- bluedino 10y ago>> from manual, to scriptable, to automated, to a fourth category I hadn't even known existed: autonomous. What's the difference between automated and autonomous?
- mikecb 10y agoI can't say it better than they did: https://landing.google.com/sre/book/chapters/automation-at-google.html https://landing.google.com/sre/book/chapters/automation-at-g...
- xapata 10y agoAutonomous does not require any human interaction, not even to trigger a script. In contrast, "automated" systems might still require a human to start the process or have a very simplistic trigger, such as based on the time of day.
- ben_jones 10y agoIf anything it proves how much mismanagement / wasted potential software organizations have had in the last 20 years. Full automation should be the natural progression of our trade but I fear most companies stall after 5ish years due to turnover, brain-drain, re-organizations, acquisitions, management incompetency etc. Google on the other hand has always had a seemingly never ending pool of resources and talent to keep pushing the barrier further. Fortunately they give a lot back to the community in the form of books, talks, and projects such as Kubernetes (a poor man's Borg). However I fear that with all things commercial it will lead to an oligarchy where companies like Google, Facebook, and Uber, are just that far ahead of the curve nobody else will ever catch up.
- kyrra 10y agoI'm a google employee, opinions are my own. The incident management tool is not Outalator. Outalator is the pager queue management tool. The incident management tool is for manually creating incidents that have much broader visibility than Outlator does. As someone who has been incident commander a few times, incidents tend to have broader impact beyond your immediate team or owned jobs.
- mikecb 10y agoThanks, my bad. Would you say the stackdriver incidents tool was influenced by SRE experiences?
- advisedwang 10y agoThe Stackdriver incidents tool existed before it Stackdriver was bought by Google. They do have similarities though.
- nodesocket 10y agoprogression of system operations from manual, to scriptable, to automated, to a fourth category I hadn't even known existed: autonomous. Completely agree. The first "eureka" moment is when you define all your infrastructure in code in something like Terraform. Magically networking, firewall rules, disks, instances, are all provisioned and dependencies calculated. It is quite a breakthrough from running CLI commands or using the web interface to allocate infrastructure. Plug: I wrote a blog post on getting started with Terraform and Google Compute Engine for those interetested https://blog.elasticbyte.net/getting-started-with-terraform-and-google-compute-engine/ https://blog.elasticbyte.net/getting-started-with-terraform-...
- asuffield 10y ago(I'm a Google SRE. My opinions are my own.) That's not what our 20% time is for, and 20% is way too small a number for that purpose. "20% time" (the way we use the term) is for personal/career growth/scratching itches. Time spent on building systems that make our service better is my primary job. Manual remediation ("toil") is something to be tracked as a dangerous antipattern that must not be allowed to take over. Toil and oncall response should be less than 20% of my time, together. At least half my time should go into engineering projects. If the level of toil is in excess of 50% of team activity then I would expect only percussive intervention to get the team out of this situation.
- mikecb 10y agoGreat comment, thanks for the clarification. Wasn't trying to say that 20% is a magic number, just that it cemented the idea for me that engineering time, and self-directed engineering time, is incredibly valuable for everyone that can be justified and should be zealously protected.
- deepakhj 10y agoHow do you track how much time you spend on manual vs projects? Is there an in house tracking tool?