4 ms·
Typo in the WaPo article - site reliabilty manager. Perhaps off-topic but are (Google) SREs the Ops-equivalent of Data Scientists in terms of expectation of cr
by blutoot 12y ago
Typo in the WaPo article - site reliabilty manager.
Perhaps off-topic but are (Google) SREs the Ops-equivalent of Data Scientists in terms of expectation of cross-domain skills? Mikey Dickerson's LinkedIn profile [0] shows he's a champ in S/W Engg + Distributed Systems + Sys-Admin + what-not. I know calling each of them a "domain" is a bit of a stretch but I couldn't come up with a better term.
[0] www.linkedin.com/in/mikeydickerson
- dekhn 12y agoThe typical SRE is a PhD in physics who realized they enjoyed hacking on the departmental cluster more than they enjoyed writing papers. Well, OK, that's a stretch but: To a one, SREs are the smartest group of people I've met, with a ton of practical knowledge about distributed systems, a strong quantitative bent, and a desire to fix problems on live systems. If you want to call that the Ops equivalent of data scientists, fine, but I prefer think of them of scientists who study the failure of distributed systems in a live environment.
- blutoot 12y agoThanks for that perspective. Do they ever find time to do side-projects? It sounds like these people work 24/7.
- dekhn 12y agoAbsolutely. They're busy but work on side projects. Depends on the person.
- eropple 12y agoI'm a devops engineer (I guess we're called SREs now too) at Localytics. We're not Google-scale, but our stuff isn't small, and a big focus of what we're doing is making sure we don't have to work 24/7. Automation, automation, automation. Handle failure states automatically and only inform humans after the system has fixed itself or if the system can't fix itself. It's pretty fun.
- blutoot 12y agoHow do you control the quality of that automation code? Doesn't that add an overhead to your operations?
- dekhn 12y agoGood testing. I got hired as a Test Engineer in SRE at Google to write automation code tests. Good coverage from unit tests, good integration tests, and good system tests (many of them using fake systems with scripted failure modes). Then finally, a fair amount of testing on "real" machines as new code gets integrated. Then make sure you have great monitoring because it will still fail at runtime and you're the poor sap on call who has to deal with it.