6 ms·
What if you encounter an nginx bug, or a kernel bug? At some point, when you reach a low enough level you will need some deeper investigation tools and finally
by kyran_adept 6y ago
What if you encounter an nginx bug, or a kernel bug? At some point, when you reach a low enough level you will need some deeper investigation tools and finally access to the machine to get enough data and fix the problem.
- paranoidrobot 6y agoThat should be the exception, rather than the rule.
- anamexis 6y agoAnd it is deeply frustrating when you run up against one of these exceptions and need to wade through some bureaucracy before you can investigate further.
- gabereiser 6y agothe point people are trying to make is that if you are at the scale where a kernel bug or an nginx bug is borking your app, it's not the developers job to go poking around the system for a fix. It's the devops/infra people's job. In my world, if you want to investigate an nginx bug... "docker run -it nginx:latest /bin/bash" and go for it... find the issue, reproduce it, then fix it in the pipeline and deploy again. You didn't touch production at all. If your debugging relies on being ON PRODUCTION, you don't suffer from the scale you need to be on there in the first place.
- jclulow 6y agoNot all bugs are sufficiently cheaply reproducible outside the environment in which they are observed. It seems silly to tie your hands behind your back when you could just inspect what the computer is doing and then fix it.
- shawnz 6y agoWhy can't your developers be "devops people"?
- paranoidrobot 6y agoI've already written a response to this elsewhere in the thread, but developers are not all equal. You can't hire twenty developers that all have the same skill/inclinations, the same interests, the same experience. That's not to say that a DevOps Engineer is some super 10x rockstar developer - no, they're going to have the same variations on skill, interests, experience, etc. It depends on your environment, but there's so much different tech once you count the entire stack, that I don't think it's reasonable to expect any one person to be an expert on all of it, or even a lot of it.
- shawnz 6y agoSure, but "not all developers can be devops engineers" doesn't necessarily imply "none of your developers should have server access".
- paranoidrobot 6y agoLets go back to the original core assertion for the thread - > > Every developer needs access to some servers for example to check the application logs. > I fundamentally disagree with this. So, Developers shouldn't be reaching for SSH access to check logs. If you're encountering problems that you can't diagnose through the existing logs, then you should probably be involving at least one other person - someone who has that production experience, who might have some additional knowledge about the problem. If, and only if you've exhausted other avenues - then reach out for SSH access. But it should be a last resort, not the first resort. Plus, anyone SSHing into production boxes should really be very familiar with how production is configured. You can do more harm than good by poking around on a production box being completely unaware that you're causing alarms and outages elsewhere because you taking a memory dump of nginx caused in-flight requests to get timeouts and so-forth. The people with that experience are generally the DevOps/Infrastructrue folks since they're the ones who deal with production all day, and are going to get the pages if something goes wrong with that.
- paranoidrobot 6y agoIt would seem to me that this is the perfect time to pull in someone with more production experience. Perhaps they can use the existing tools to pull logs, or analyse it in some way. Maybe they've seen it before and already know the fix. Giving everyone production SSH experience is, in my experience, a way to run into all sorts of weirdness, not to mention endless frustration. In a modern automated infrastructure, that box is likely a container running on a virtual machine that's ephemeral and can (and probably will) go away at any moment based on any number reasons - maybe CD kicked off a new deployment, or maybe the load changed and the instance was selected for scale-down, or maybe our spot bid for that AZ isn't sufficient for keeping the instance around, maybe you being SSHed in and poking around impacted the health-check, and so it's being killed for not performing right. Theres many other problems, too - lots of applications are built in some way that there's simply no other way than secrets (passwords, api tokens, keys) to reach other systems, particularly third party systems. So production boxes have production secrets, which you probably don't want to share with everyone. Giving everyone SSH access so they can, in theory, take nginx/kernel dumps as needed tends to imply giving superuser rights, which means they can do whatever they like. So, yes, pull in someone else - find some way to try and reproduce the problem NOT on production, if that fails, perhaps there's a way to grab enough detail or pull additional logs or network captures to identify the issue. If that fails, well okay, lets SSH in - but we need to coordinate that to ensure that instance does't go away, and doesn't impact production while you do it.
- hnlmorg 6y agoThen that would be a sysadmin / DevOps responsibility rather than your application developers. Or at least you'd have those guys involved with the investigation. But honestly, how often is a web applications bug due to a kernel bug?
- Aeolun 6y agoYeah, ever since devops became a thing they’ve placed themselves in this position where they’re somehow better than the application developers, even though until a few years ago those same developers were doing the exact same things. I swear, sysadmins were annoying as an application developer, but devops is something else. People with one year of actual work experience get hired as devops, and have all the privileges I would need to fix their mistakes, but I can’t, because I’m an ‘application developer’. So instead you end up teaching them how to do their job. I’m not salty at all.
- xref 6y agoAh the classic ivory tower argument where some “other class” of engineers are universally inept, but not “my class!” You can write the same screed full of generalizations from the perspective of any job title: a devops person would lament the fresh-out-of-bootcamp “application developers” who have no idea how systems work together so write SQL queries that retrieve a million rows, one at a time. “Works on my local!”
- hinkley 6y agoPretty sure GP was bristling at the reverse happening. We must keep the developers from screwing up the important computers. Saying the emperor has no clothes is not white tower thinking,
- Aeolun 6y agoI completely agree. Access to those things should be given to those qualified to work with them, not based on an arbitrary role designation.
- 6y ago
- athms 6y ago>What if you encounter an nginx bug, or a kernel bug? That is the responsibility of system administrators. Application developers have no business on a production machine. If your sysadmins don't have the technical skills to diagnose these problems, they are incompetent and must be replaced.
- Aeolun 6y ago> If your sysadmins don't have the technical skills to diagnose these problems, they are incompetent and must be replaced. The actual result of this is that the sysadmins are not replaced, and the application developers end up in an emergency conference call at 3am to tell the sysadmins which buttons to click on the production environment, since they’re not allowed access themselves.
- toomuchtodo 6y agoIf developers aren’t exposed to the deficiencies in their systems, they have no incentive to reduce SRE pages and triage. Build resilient code with quality documentation and you don’t have to attend a 3am conf call. DevOps is not a role or role segregation, it’s about aligning incentives and outcomes across functions in an org (hopefully through collaboration, tooling, and knowledge transfer). The caveat is that if your org is fundamentally broken, none of the above applies or works and it’s all lipstick on a pig.
- athms 6y agoI spent several years at a large multinational cloud provider that gave developers and QA access to production systems and customer PII. That all changed after the company was bought by SAP and operations were integrated. I am amazed that engineers think this is acceptable. It is bad business practice, compromises security, and illegal in some jurisdictions.
- rhizome 6y agoHere's a funny thing: as of two days after this post was created, pairing hasn't been mentioned once in the entire thread. If this thread is any indication, maybe it's developers who have an incomplete understanding of DevOps.