6 ms·
I've been each of them over 5 years and been involved in a huge enterprise. In my experience: - DevOps: Hired to do everything not involved in the feature deve
by CSDude 5y ago
I've been each of them over 5 years and been involved in a huge enterprise. In my experience:
- DevOps: Hired to do everything not involved in the feature development of main business. Range varies from Terraform to maintaining Jira, GSuite, Jenkins, ETL, BizOps. Maintains scripts and hacks everywhere.
- SRE: Orders teams to define SLOs, focuses only after incident procedures, not before. Insists on creating dashboards and putting them on TVs (when we had offices). Most eventually become vendor fans.
- Platform: We made this, we know better than you, you cannot use anything else. If you need new thing/improvement, create a ticket we will get back to you in 6 months to say we cannot do it due to other company wide very important initatives.
I love not being a regular developer and work on platform side, but its a very irregular space and very hard to hire.
- royjacobs 5y agoAs a platform engineer, that's quite a cynical take you have there. What I've found works best is to write tools that take away complexity so that teams don't all need to reinvent the wheel. Then make it so well-documented and easy to use that teams don't need to be _forced_ to use the platform tooling, they'll just use it because it's the most convenient thing to do.
- CSDude 5y agoI mostly agree with you, an average dev does not know and do not care how things are done. But if platform team or any tooling is blocking you from delivering your features frequently, it becomes an issue. They are either under-staffed, or generalizing in favor of greater good comes at a cost of dropping your specific requirements.
- pydry 5y agoThat's a wonderful idea in theory but I've never seen it work out that way in practice. I think it's just genuinely too difficult for most companies to do well and it's too easy to fall into the same old well trodden traps. Documentation is something everybody professes to keep up to date, clear and complete, for instance, but IME it never truly is.
- geofft 5y agoThey appear to all be cynical takes. :) I think one way of reading the comment above is that it's the effect of business incentives / Conway's law on various approaches to staffing the non-developer parts of your organization: - DevOps is incentivized to be an extension of the dev team and get stuff into production, i.e., in the average company hiring "DevOps," the engineers operating production report up into existing dev management. Therefore you get "scripts and hacks" because you're never given the mandate to do anything better, and you're evaluated on how much you can manage to ship things to prod without building separate specialties, not how well you ship it to prod. - SRE is incentivized to demonstrate good SRE practices and their value to the business. Incident response is a lot more visible than incident prevention. Picking an observabiity vendor who can get you nice graphs is a lot more visible than spending months making the right graphs. You're evaluated on how much you improve "reliability," which means you need to figure out how to measure it (and display it) in the first place, but it's not clear your measurement is what the broader business would measure. - Platform engineering is incentivized to build up a well-defined platform and have their own products/services that are used internally. They are a development team of their own. So their value is most visible when those products exist and are widely used, and a little more cynically, when those products aren't unmodified vendor or OSS products. You're evaluated on the number of people using your product and how infrequently the rest of the company says "I'm just going to run this myself, give me an AWS account." Any of these disciplines can be done well with strong management support, by which I mean when management actually cares about the long-term problem of running production well and can make informed decisions about how to staff that problem. (For myself, for what it's worth, I'm in a platform engineering organization but my current role is 80% to make it easier for people to use OSS and to reduce our deltas and 20% incident prevention.) But when incentives get disconnected from the business and the inertia is to keep delivering what people thought you should have delivered in the past, the disciplines break down in different ways.
- midrus 5y agoIn my experience, the end result is having a team of kubernetes fanatics reinventing the wheels themselves and forcing everyone to use their custom ego filling tools instead of paying $100 per month to an already existing solution. Most developers just need something like heroku, tsuru, appengine, or openstack. But because kubefans want to have fun too and try to be as cool as Google or Facebook we end up having to deal with tons of custom Golang tooling with fancy names, half assed documentation and grumpy looks when you ask for help because you don't understand the 38 steps guide to deploy a service.
- sascha_sl 5y agoWe're not here to compete for most cynical take. $100 is not really a price point for any of the things you mentioned, and the way PaaS charge is often restrictive of development model (want microservices? we charge by the service. want a monolith? sorry, we only have instances up to this size. need to scale in smaller increments than "one machine"? too bad.) When I helped build my first platform, it was a migration from Beanstalk to Kubernetes (1.8 at the time). Not only did it save 70%+ costs, it allowed developers to cut services as large as they wanted, and scale them as fast or slow as they wanted. Metrics and logging were nothing they had to worry about on the infra layer either. I wouldn't build these things again, OAM / KubeVela is a nice compromise if you want one. Also, I wouldn't give a dev direct access to Openstack the same way I woudn't give them access to Kubernetes itself. They have too many footguns and too much complexity, either you end up with half-baked deployments or half the teams time spent on doing things the right way.
- midrus 5y agoOf course it is not $100, but it is a lot less than the salaries of a team of infrastructure engineers most of the times
- sascha_sl 5y agoAnecdotally, when my entire team started leaving over the span of 5 months or so after a new CEO attempted to force a move to every service provider he had previously worked with (GCP over AwS, Cloudflare over Akamai, MS Teams over Slack), the non-cloud costs for the platform jumped to about 6 times the teams combined salary when they decided to hire specialist contractors, and even they had a hard time matching the raw performance of that cluster. The typical case of feature teams not being allowed to refactor / reduce technical debt was extreme. Some frontend requests would result in hundreds of backend requests for no good reason. We've always been ignored by management, so cluster was on the verge of coming apart at the (mostly "spending half the cpu time in conntrack") seams for years. So close, it'd sometimes just die for a minute, probably because someone decided to use HTTP way too liberally (network is free, right?). I had a pcap file from a sample of 10 machines, and a 30 second sample was so large, it was impossible to analyze with any real-time tools. No managed Kubernetes could match this out of the box. Adding a mesh to do more directed requests for this to be fixed failed even when we tried it. As far as I know they're still using the platform as is (in maintenance mode, with the contractor assisting) and fired that CEO pretty soon after. Is this a bit of a special case? Sure! Does it still happen a lot, no matter your infrastructure? Count on it.
- eddieroger 5y ago> What I've found works best is to write tools that take away complexity so that teams don't all need to reinvent the wheel. The problem is complexity is variable. For developers who never really learn how their code actually executes, or haven't had to be responsible for tuning and such, this is great, but for those of us who want to know where the rubber meets the road, these tools are kid-gloves. How can we fully optimize and support our code if we can't see the compute on which it runs? What other "benefits" will the system bring that we have to account for? > Then make it so well-documented and easy to use that teams don't need to be _forced_ to use the platform tooling, they'll just use it because it's the most convenient thing to do. Very optimistic. In my experience with this kind of thing with homegrown platforms, documentation was a distant last, because it always is, and we weren't left with other choices anyway. I would love to pick the best tool because it is convenient, but that, too, is variable. Maybe I know Kubernetes inside and out - that's plenty convenient for me, then, and will be better documented and discussed than any homegrown platform because the world can contribute to it.
- crypt1d 5y agoYou are mistaking your personal experience with various companies and positions to the actual industry standard. I've worked on the same roles as you did and my experience is completely different and very positive.
- CSDude 5y agoI explicitly mentioned it's my personal experience. I'm normally very happy with what I do. Not just the enterprise environment I'm in.
- TurboHaskal 5y agoOne should not disclose that their comment is based on their personal experience when it is obvious this was the case. My personal experience agrees 100% with what he has said :)
- quadrifoliate 5y ago> - Platform: We made this, we know better than you, you cannot use anything else. If you need new thing/improvement, create a ticket we will get back to you in 6 months to say we cannot do it due to other company wide very important initatives. I think it is due to cynical takes like this that the industry is evolving towards “Use whatever you want, so long as you are open to being on-call for it yourself. Or you can use this Platform maintained by other people.” Nothing sucks as much as being woken up by pages because some new, shiny, and untested piece of technology was used in production and broke in an unknown way.
- pram 5y agoThen the team who made their own platform hire DevOps and SRE to janitor what they created because it's spiraling out of control, and now you have another platform! It's the circle of IT life.
- CSDude 5y agoWhat's cynical about it? If the platform team is on the way it becomes a problem. And going to be on-call is way to go if you want to have actual ownership, or it becomes to same dev and ops people split. There is not an exact line. Platform team can easily abstract (i.e) Kubernetes deployment infrastructure, but not everything else that you might be needing.
- debarshri 5y agoI have seen organisation where platform teams build tools, frameworks and nobody uses them. Few places where it makes sense is when an organisation is large and has corporate and regulatory policies that applications and people have to adhere to. As platform team member, you really have to know what to build, doing interviews, figuring out what developers and SREs really want, filter through what they are asking for, control through not-invented here syndrome, make build-vs-buy decision, create the community around the product you build within the organisation etc. You cannot just start building a PaaS (most common thing I have seen them build). It is actually quite complex if done right. Having said that, lot of organisation the definition is quite vague, political too and sometimes the boundaries are not clear.
- wayoutthere 5y agoI do a lot of this kind of work on the consulting side, and the interview process is critical to establishing buy-in from the essential stakeholders. You have to understand the pain points because without that, you won’t build the right thing for the way the company operates. The most common hurdle is simply that most developers buried in enterprise software dev have no familiarity with the cloud or foundational DevOps tech like infra as code. So we typically end up building some form of deployment mechanism cobbled together with packer and terraform, but the real value in bringing in outside help is that we handle training and change management necessary to actually drive adoption. “Build it and they will come” is a terrible design philosophy. You can’t just stand up a platform team and expect it to go well: a more transformational mindset inclusive of org design and strategic outlook is required, and this type of DevOps / SRE transformation is no less of a radical change than HR or finance transformation programs.
- halfmatthalfcat 5y agoOur systems team is veering into the first sentence hard. Tried to make a whole "one size fits all" CI/CD CLI tool that is inflexible and obfuscates the process. Nobody asked for it and everyone is content rolling their own pipelines by hand, yet they've sunk a ton of time into it.
- 5y ago
- lamontcg 5y agoHow did SRE become such a non-skilled job almost like a CISSP auditor is for security? Originally it was supposed to be a SWE who could also do operations, but there seems to have been a quick race to the bottom to remove all the SWE tasks and turn them into pure operations policy wonks.
- dilyevsky 5y agoIt was always like that outside of select few like Google. Most companies’ infra is not complex enough to warrant anything beyond sysadmin skills. Fwiw google sres also had two distinct subroles within sre one of which is not software focused
- debarshri 5y agoSRE is a highly skilled job. It is actually difficult. Doing a root cause analysis or investigating an infrastructure blackhole is a very rare skill. You have to be aware of the end-to-end landscape, or atleast communicate well with the stakeholders who own the domain and resolve issues, manage expectations, huge pressure dealing with time critical issues. Having seen all of roles, being SRE was the only time when a slack message sounds would trigger anxiety in me.
- lamontcg 5y agoI mean I used to be a tier 4 oncall at Amazon and had global root access, managed the configuration management infrastructure and had enable access on all the networking equipment. I'm criticizing the SRE role from a background of having a very high level of understanding of difficult operational roles. The selling point to me 10 years ago was that it would break down the wall so that people like me could also have access to the source code and be SWEs (instead of having to strace a black box all the time), but it seems like the job has become significantly dumbed down from the systems side, and SREs tend to grow out of the software teams and are the people responsible for the pager, metrics, SLOs, hacking up turing computable YAML, memorizing the whole CNCF landscape, and yelling at Kubernetes.
- 5y ago
- noknownsender 5y ago>Insists on creating dashboards and putting them on TVs You didn't have to call me out like that.