6 ms·
Anyone who has ever seen the deployment diagram of Google's ad serving will vouch that Google simply cannot exist without great SREs. If you like both dev ops a
by general_ai 10y ago
Anyone who has ever seen the deployment diagram of Google's ad serving will vouch that Google simply cannot exist without great SREs. If you like both dev ops and software engineering, and have found that your affinity to dev ops makes you a black sheep, I encourage you to apply to an SRE position at Google. I can state unequivocally, that SRE's are held in great regard at Google, and they receive a tremendous amount of respect. This is helped somewhat by the fact that you have to actually earn their support. Until your service is considered maintainable and observable enough to not cause pain, you'll be doing your own DevOps. It's only when you pass the PRR (production readiness review) that you _might_ get _some_ SRE help.
Disclosure: I'm a former Google employee
- packetslave 10y agoyep, and if your service gets significantly less reliable over time after SRE takes it on... well, either that will get fixed or you'll be taking the pager back until it does.
- tehlike 10y agoGoogle SREs are pretty good, but there is disconnect between product and sre (somewhat understandably). SREs will always take steps to mitigate issue. rolling back binaries, mlts, gdps, you name it. Sometimes, it's actually easier to fix the problem, but SRE doesn't know it. It's understandable, they manage many many projects, and they can't know all the products. It's just something i found very interesting. Obvious PS: Google employee.
- geofft 10y agoWhat's the right role for someone who wants to deeply know some products enough to fix the code the right way, but doesn't want to be a dev? (Be a dev somewhere where maintenance is as valued as creation?)
- rumcajz 10y agoI think SRE is what you are looking for. They typically know the product pretty well. The reason why they rollback rather than fix the bug immediately is that they want the outage to by fixed in minutes. Even if you spotted the bug immediately you would not have enough time to do the build, let alone run the tests.
- deleted 10y ago[deleted]
- willemmali 10y agoI think that's about right. I think I just started doing work that sounds very much like SRE work to me: I'm building a CI pipeline, E2E tests and "Dockerizing" an existing Java-based project management product (currently only deployed as SAAS, but on-site deployment is in the backlog). I'm trying to fully automate the testing side of the product, while making the process transparent enough to be amenable to manual intervention/quick tweaking. After that, I'm hoping to move to automating the deployments, putting the server behind a load balancer, rollbacks, backup testing, all that good stuff that makes sure things only break where it can't hurt. Luckily the product is already pretty stable with the current dev/dogfooding-as-staging/prod model. It's the most enjoyable work I've had so far. I think it mostly boils down to: * I have clearly defined tasks, which I mostly plan in/negotiate with the product owner myself, so I have a large share of "ownership" of the dev/QA infrastructure improvements * I work fully remotely and part-time, which gives me plenty of free time to socialize and decompress (we mainly communicate via Slack). I also have the option of working more hours, but I already doubled my * I'm not currently on the critical path, so work feels low-stress * I don't have to deal with under-defined business logic and product owners that do not want to commit to specifying (the product owner has transitioned from building the Java software to managing and subcontracting it, so is very knowledgeable about the product, and besides he's a great guy) * I'm learning the tooling around the product through automating its development, testing and deployment (vs. learning it through adding crufty new features to it in a completely un-repeatable way, I'm looking at you never-again crashing Visual Studio Community and randomly-failing-builds Xamarin Forms).
- jpatokal 10y agoTechnical solutions engineering (TSE) may fit the bill, especially for more mature products. Think support on steroids, where you're empowered to fix the customer's problem. Obligatory disclaimer: I'm one, and we're hiring ;)
- tehlike 10y agoWhy do you not want to be a dev?
- packetslave 10y agoWell, I think it's less of a disconnect than a difference in priority. SRE's first priority is "stop the bleeding" -- take whatever immediate action you can to stop users from being hurt. That might mean rolling back the binary, reverting a data push, draining away from a broken cell, whatever. When you're serving thousands or millions of QPS, time is of the essence. That being said, SRE does want to ultimately fix the problem (otherwise it's just going to page again, right?). But if that means tracking down a wrong config flag, cherry-picking a fix into a new release, etc. -- those are all things that can be done AFTER the bleeding is stopped. Source: I'm an SRE
- deleted 10y ago[deleted]
- deleted 10y ago[deleted]
- tehlike 10y agoOne of the cases i was involved was when the issue was not found after 30 minutes, after sre rolled back most of the systems. Reproducing the issue resulted in an immediate fix by the swe. Again, i understand why it is the way it is, it is just really interesting to see how specialized each engineer is in the grand scheme of things.
- general_ai 10y agoImmediate fix by SWE can only be released after it is tested and canaried, so it's not really "immediate" most of the time.
- sokoloff 10y agoWe run factories and if we had a bad deploy bring a factory down, it's not going to get "more broken", so we can push a fix-attempt change live as soon as it's ready. Abstractly, we got pushback from QA about this policy. After we had gathered a couple of concrete examples, it was clear that QA-as-gatekeeper when the factory is already in the worst possible state wasn't valuable. We do mandate the normal reviews but allow them after deployment. (You can imagine the conversations with the auditors about this as well, so we had to carefully document that this was our process and made the auditors audit our conformance to the process not to their own preconception of what it should be.)
- brobinson 10y agoGot a link to the mentioned ad serving diagram?
- general_ai 10y agoGet hired by Google and you'll be able to check it out. It bears a close resemblance to a Rube Goldberg machine.
- manglav 10y agoHow do you break into that? I'm a mid level software engineer looking to get into devops roles, and have been told by an ex-youtuber that I have a real knack for it. When I applied, I was rejected for not having enough experience. I find tooling and devops extremely fun, but I'm not quite sure how to develop my talent. Do you have any advice?
- general_ai 10y agoIt's mostly a matter of chance, TBH. You miss 100% of shots you don't take. It took me two attempts to get in, first time I was being stupid during phone screen so it went no further. I was a SWE, but it's not generally a problem to switch. As a SWE you can even switch temporarily for 6 months to really understand how systems work, and what not to do when building them.
- deleted 10y ago[deleted]
- zippergz 10y agoMy problem is that I enjoy many aspects of SRE work, but I absolutely despise being on call. I've since transitioned into onto a different career track, but I have long wished to find some way to use my combination of unix sysadmin and software engineering skills without ever having to be on call. In the companies I worked in (including Google), I never really found that.
- deleted 10y ago[deleted]
- general_ai 10y agoEven SWEs are oncall. Not as much as SREs, but it's really unavoidable. Someone needs to be there if shit hits the fan.
- zippergz 10y agoDepends on the project. As a swe I got paged zero times.
- general_ai 10y agoFor anything that actually receives significant traffic, runs in a lot of cells, and is under active development, Telebot will call you pretty regularly when you're oncall. But hey, congrats for finding a place you like, that's exactly how things are supposed to work there.
- twright0 10y agoYou can be on call without ever getting paged. As long as the expectation exists that you're available, it doesn't matter whether you actually get paged or not - you're already shifting your plans so you stay within reach of a laptop, with cell signal, etc. Being paged is the least bad part of being on call. The restrictions on your life outside of work are what really grate.
- CydeWeys 10y ago
- deleted 10y ago[deleted]
- fh973 10y agoThe grandiose production readiness review, didn't hear that term for a while. In 5 years as a Google SWE, having transitioned three systems to SRE, nobody could ever tell me what it really is. In each hand over, someone just came up with a random checklist. I am really curious: is there really a disciplined/formalized/.. PRR process in some parts of Google? Has anyone ever seen it?
- general_ai 10y agoYou basically do the shit SREs tell you to do if you want support. I took a service through a PRR, and while it wasn't 100% formal, my SRE peers were able to request improvements to monitoring, fault tolerance and release process, so it worked well in the end. Other than the launch checklist (is it still called Ariane?), Google has few truly formal processes in general. People converge on what works for them.
- bluecmd 10y agoIt's up to the SREs taking over the service. A storage SRE PRR is different from an Ads SRE PRR.
- dannypgh 10y agoI've seen several formalized PRR (in a meeting about improving cross-team PRRs for pipeline systems) forms. If you were getting a "random" checklist, the SREs were doing more total work- the work of writing a checklist tailored for your service. Lots of questions don't apply to lots of services, so the standardized forms tend to help whoever is giving them out at the expense of whomever is reading them. Very common questions get at basic things- what are the pain points you've encountered running the system? What monitoring systems are you using? What are known failure modes where monitoring is silent? Have any agreements on availability or latency/performance been reached with users? What is the process to qualify, push, and rollback changes? What's the impact to the user / to the company if everything goes and stays down? What are your runtime dependencies and how do you behave if they fail? Provide a review of recent monitoring alerts[1]? Most of the value in most PRR checklists really just get at the above- sometimes the answer really is "we don't know" or it is incomplete (especially re: runtime dependencies) so follow up questions can make discovery easier. [1] often the SREs can figure this out and will look them up without even asking a question. Lots of formalized processes ask people to list what's needed to do this anyway (e.g. list alert queues, mailing lists that receive alerts, etc.).
- pmoriarty 10y agoSome of us have ethical problems with working for a company whose business model is centered around spying on people.