10 ms·
Why should I do production support?
- MattGaiser 6y agoPeople seem to get stuck in support work. That would be my aversion to doing too much of it. At some of my prior workplaces, there have been people who have been so good at support that they never got assigned to do new development work. Engineers far more experienced and senior than I was (or currently am) were dealing with trivial issues as they were good at it while I got the nice greenfield project. They had to quit to get out of support.
- deleted 6y ago[deleted]
- x87678r 6y agoThis has been a huge problem for me. I've always loved working with customers and dealing with real problems under pressure. So have enjoyed my time on support rota - but quickly L1 guys learn I'm better at support that other team members so everyone contacted me directly. The grumpy devs who pretend they dont know anything about the live system get interesting projects, no interruptions and a much better resume. I've learned to say no. Its also a good reason to move teams as you're not the super experienced guy who has to fix the urgent gnarly support issues.
- runawaybottle 6y agoSome people have a switch then turns them into bug fixing demons. I know I get a rush out of fixing prod issues. People notice that, especially when you just jump into foreign code that you never touched and come out with a fix within an hour or two.
- jasonlotito 6y agoOf course, if they were the ones who originally built those products, they should support them. Why should people that create software that requires so much support they quit over it be entrusted with yet another greenfield project without fixing the stuff they built in the first place?
- greesil 6y agoAt some large tech companies, production support as a software engineer does not seem to be a path to promotion, unless you are a junior level engineer. And yet, solving some of the bugs encountered in production environments, especially with a heterogeneous set of users, requires expert level knowledge of a particular software library. It is probably a great way to coast, so I hear.
- goatinaboat 6y agoIt is probably a great way to coast, so I hear. Or to stagnate, depending on how you look at it
- aprinsen 6y agoTeams I've worked on have not had dedicated support staff, but rather engineers rotate 24 hour support duties on a weekly basis. I have always had mixed feelings about "on call". I dread my turn on the rotation because the imminent threat of a prod issue has a psychological impact on my entire week, even off hours, and usually for a day or two after. If everybody on the team feels that way, maybe it can act as a forcing function for product quality. I've seen this work on teams that already cultivate a strong sense of ownership. On the flip side, it really stresses me out, and I sometimes resent that I'm not getting paid overtime for 24hr on call days. Maybe that's just baked into an engineer's salary these days, though...
- MattGaiser 6y agoThe engineers at the organization I just departed (at least the ones in the support rotation which did not include me) got paid in both money and extra time off for their time spent on support tasks.
- kevinmchugh 6y agoYeah, I had a job where only certain teams had on call rotations. Anyone in those rotations got an extra half day off per week on call. I've also seen folks spotted extra time off for really gnarly oncall shifts. Folks should push to have such accommodations standardized.
- user5994461 6y agoA half day off for doing 7 full days of extra work. It's peanuts. Better not do the rotation and have your 2 days of week end. (I assume they got zero extra pay and got called occasionally).
- kevinmchugh 6y agoOn call shifts were very quiet and typically didn't involve any pages or interruptions. It was certainly not even a full day of extra work typically.
- woutr_be 6y agoI work in finance, where we have clearly defined processes for production support, simply because developers don’t have access to production environments. However, production support teams don’t have a real understanding of our application and how it’s build. So most of the times you have engineers on call with production support, telling them how to debug the problem and come up with relevant logs. It’s incredibly infuriating and time consuming, and I absolutely hate doing it this way. 90% of the time you also get incredibly vague bug reports with irrelevant logs, and a description of what they think the problem is. Most of the time you need to spend another day finding correct logs and somehow debugging it. Most teams log every single request with all parameters and payloads because they can just replicate the problem locally instead of relying on production support. We’ve long advocated for either having dedicated support or have engineers on some sort of schedule that can do support.
- pawelmi 6y agoI'm curious, why can't people that created the system take part in production support. You've mentioned finance, that I presume require high level of security, but at the same time there are also people on the other side, just not knowing the system first hand and perhaps having skills different from knowing how to debug software. They can see all the data and in theory modify system behavior, eg modify/install any binary. Why are thrall developers less trusted, is it some kind of logic or regulation or just "the way it has always been done" in finance?
- goatinaboat 6y agoI'm curious, why can't people that created the system take part in production support. You've mentioned finance, that I presume require high level of security, but at the same time there are also people on the other side, just not knowing the system first hand and perhaps having skills different from knowing how to debug software They can, they just can’t have direct access to live systems due to separation of duties. But there are methods for dealing with this, like centralised logging so a developer never needs to see the original log file on the problematic box.
- ocdtrekkie 6y agoIf you aren't doing production support, you don't actually know your product. You aren't connected to the pain points your users experience and you miss what is, to your support team and your customers, the glaringly obvious. I would argue all developers should be required to do some support work.
- axaxs 6y agoI largely agree. My company has dedicated CS, a second tier 'triage', and well...me. I always prefer if customers just email me directly. CS is a frustrating, cost saving measure. And often I'll get tickets, 3 weeks later, like 'customer said they have an issue.' What's the point again?
- ed25519FUUU 6y agoHaving skin in the game will make you a better engineer. You’ll get better at the non-sexy things: monitoring, alerting, testing, etc. You better believe a person who is pages for software at 3:00 am has an incentive to make that software more reliable.
- lostdog 6y agoBut the directors and VP's aren't getting paged at 3am, so the incentives of the organization still go to lower quality software.
- johnbellone 6y agoIf the directors and VPs aren’t getting paged at 3am you need better leadership. The first thing that I did was subscribe to all outages. And I let the operations center to call me anytime if they need help resolving a production issue.
- ocdtrekkie 6y agoI'm not even talking 24/7 pager work. But just tackling some support tickets as part of your job so you see where people are having issues with your product. Too often I see BigCorp development teams seeming blatantly oblivious to where their pain points are, and it's because they aren't forcing their developers to do support. They're pushing code, but they aren't pushing code that solves real problems for people.
- AkshatM 6y agoIt sounds very much like the poster is describing an on-call rotation rather than "production support", which is a very different thing altogether. Production support is customer support: responding to chat messages or communications from users. An on-call rotation, on the other hand, involves responding to production incidents and mounting a proper incident response. The Google SRE workbook has a great chapter on the subject: https://landing.google.com/sre/workbook/chapters/on-call/ https://landing.google.com/sre/workbook/chapters/on-call/
- johnbellone 6y agoI don’t know why you’re being downvoted; these are two entirely different roles.
- crossroadsguy 6y agoYou don't get anything out of burning out. You just burn out. And that time is not coming back. This post seems like an apologist talking. Production support alone is not that much of a problem. What the author skipped (conveniently? or forgot to mention?) is - it's really the "on call" phenomenon that's the problem. The "typical" on-call - where when you are on-call you are magically on-call 24x7. Yes, during your sleeping hours as well; as if that's less important and the company can avoid spending money to hire dedicated support for those hours and instead make you suffer (yes, it's just that - there's no other name for it like "satisfaction", "learning", "growing" or any of those buzzwords). You want engineers to do production support? Well, let them do it during normal office hours and only few times a month. Or heck, let them do it for weeks but let them punch in and punch out normal office hours. Let them choose to do only one half of the day and have someone else willing to do the another half. There's no excuse for burning out engineers (esp. unsuspecting youngsters) by pushing them into ungodly hours of work ruining their health among other things while trying to constantly tell them - "do you even realise what a service to humanity you are doing!". It's just exploitation.
- dopylitty 6y agoThe point about on-call is really critical and really on-point. If a company thinks an application is important enough to run 24x7 then it should staff for 24x7 support. Stealing wages from workers by expecting them to be available 24x7 (on-call) is an absolute abuse. It also leads to burn out, poor performance during the day (how is a dev's development ability when they were up at 2:30am on an incident call the night before?), and clouded thinking causing mistakes or impacting recovery time during incidents.
- HenryBemis 6y agoThe companies do this for the money. And the people what work in those companies have no real sense of the risks, or they just care more for the numbers and want to roll the dice. I would not trust someone who I just woke up at 2am to do something. He/she is mid-sleep. They will be prone to errors, they will be super tired, and I just ruined their next 1.5 days that it will take them to recover from that. This is not a job where you live boxes where intellect is not needed as much, (strength and stamina will also be affected by a mid-night alarm). You want your folks to be 100% on par, otherwise they may make things worse.
- bcbrown 6y agoThis seems very telling: > I no longer work at Gojek
- MattGaiser 6y agoWhy? Engineers switch jobs all the time.
- momokoko 6y ago> Engineers switch jobs all the time. This also seems very telling
- MattGaiser 6y agoRaises are nice?
- hyko 6y ago...because your company wants you to work two jobs for the price of one?
- efitz 6y agoI really like the author's first point, that one of his learnings was empathy for customers. Being on support for a product you didn't develop really calibrates you to understand what a product needs in order to be supportable, and where customers have problems with products. My time years ago in support was invaluable in shaping me as an engineer; I regularly push back against features that I know will be difficult to support or difficult for customers to understand.
- seanwilson 6y agoAnother pro for consultancy work: you can chose contracts that that don't require you to be on call for production problems so this isn't forced upon you. I'm not saying you wouldn't learn from working on production, but whether it's worth the stress is another question. In terms of software development, it's hard to think of a worse feeling than when you do a production deploy, you hit refresh on the website or whatever it is, and it shows a fatal error, then there's a mad scramble to roll back the change and figure out quickly what went wrong before the consequences grow too great. Most of the time bosses + coworkers aren't that understanding about it either and get into finger-pointing.
- user5994461 6y agoConsultants charge 200% for out of hours work, if not more. I don't think they mind working extra hours. They're never offered extra work though. Companies are always willing to wait for Monday when they are asked to put money on the table.
- sdevonoes 6y agoI like to write software but I don't want to be on-call if the software I wrote breaks at 3am in the morning. I do take my job with professionalism, I do write tests for most of it (not 100% coverage, but 100% coverage of the critical parts), I do monitoring (and answer and fix alerts if they happen during working hours) and I don't deploy on Fridays (and don't allow people to deploy on Fridays). My code will crash sooner or later. I already know that. I don't write 100% bug-free code. But I cannot accept to give 100% of my time one week per month or so to a company in exchange for money. I just don't understand why people can't understand that I can be a professional only during 8 hours per day, but not more.
- peterwwillis 6y agoThis attitude almost always turns into the following: On-call: "Hey devs, I'm being woken up at 3AM because your app sucks. Please fix it." Devs: "Sure, no problem." 4 months go by On-call: "These alerts are still coming in at 3AM. Did you fix the issue?" Dev: "We have a lot of work, we can't dedicate all our time to some minor problems, we have a deadline." Next week, Devs are put on-call. The alerts are fixed in two weeks. Site reliability goes up. Apps suddenly become more resilient to failure. Honestly, the whole attitude of not wanting to work more than 8 hours is privilege. Most of the rest of the world works long hours. As a dev, you get a good salary and a job you don't have to break your body to do. The least you can do is be completely responsible for your own code. And it helps you as an engineer. Like the article points out, it creates empathy for the users and product support engineers, it helps you improve architecture and app design, and it helps you understand different failure domains. You won't learn all that on your own time, especially without the scale of production.
- robmsmt 6y agoenter devops
- cabraca 6y agoPutting the blame on the devs is to easy. There are a bunch of management layers between the on-call team and the devs in your example. If management does not prioritize those alerts, its not the fault of the devs and its the wrong to put punish the devs for it. > Honestly, the whole attitude of not wanting to work more than 8 hours is privilege. Most of the rest of the world works long hours. As a dev, you get a good salary and a job you don't have to break your body to do. The least you can do is be completely responsible for your own code. unless i signed a contract that states i will do on-call, i'm not gonna do on-call. I doesn't matter how long the rest of the world works.
- hermitcrab 6y agoI am an independent developer living off 3 software products I created and sell. I have done all my own support over the last 15 years. While it can be frustrating (especially for B2C) it is also my superpower, as it gives me some much more insight into how I can improve my products. I only do support via email (not phone or chat).
- jake_morrison 6y agoA lot of the pain around production support is easily solved by having staff in multiple time zones. A good structure is to have first line support be relatively generic ops people. They can handle problems related to infrastructure, e.g. hardware failures, network problems, or issues that can be handled by adding resources. The deployment process should be consistent enough across applications that they can e.g. roll back to a previous release. This covers the majority of production problems. After that, it's time to bring in someone who understands the details of how the application works. If the dev team is geographically distributed, then someone is available during working hours. Otherwise, we have to get someone out of bed. If the dev team has done their job right, this should be a rare occasion. Making the dev team fully responsible for the reliability of the application means that they are motivated to make it reliable. Otherwise there is a tendency to have an underclass of ops people who get abused. A fundamental mindset here is taking responsibility for the user experience, including reliability. If this is not owned by the product development team, then who?
- mtberatwork 6y agoConvincing folks at the top of food chain that more staff is needed is one of the most difficult things to do.
- brailsafe 6y agoI feel like the author is referring more specifically to troubleshooting issues when production goes down, which I'm fine with if I have no other people asking me for updates. But, I burnt out at my lost job trying to do support, because there was individual support baked into every contract for our SDK and not remotely enough people to handle it all. I was hired as a software developer, not a customer support person, and they are not the same thing. It's unfortunate, because it was my first and highest paying gig after realizing that I also have ADHD, and it was a good company. Thing is, if I have a problem to solve or task to complete, I'm not going to think about how long it's been since I replied to whoever about their pet issue. I'm just going to zone in on my thing, and if the guy next to me doesn't break me out of it by cracking eggs on the desk and burping, then I'll stay on that thread till it's done. That's how my brain works, and expecting otherwise is naive. Anyway, this constant context switching and battling my apparent insufficiency killed my spirit for the work and I turned into a blob of productivity. It's as stupid as expecting me to cook while programming, because either the food, myself, or the code will get burnt.
- g051051 6y agoOur company is trying to move to "devops" and having dev teams on pager duty, doing manual reporting, etc. They seem surprised at the amount of pushback from the devs.
- comeonseriously 6y agoOn the one hand, a lot of devs see production support as beneath them. On the other I think managers seem to thing they'll get a net productivity increase doing this, but there's not. Development gets much harder when you have to context shift several times per day. That being said, I do feel devs should do prod support. It gives them a better feel for their apps they build and how they're used and where the customer pain points are. But, I also feel there should be layers of support below the devs; devs should only get the ticket when they're the only ones left who can figure it out.
- g051051 6y agoIt's not "beneath" them, but it's an entirely different skill set. If you want me involved, approach me like a partner and we can work together, don't try to suddenly give me a job I can't do.
- jp0d 6y agoIt depends on what and how much one is learning from it. I once interviewed a guy for an ETL developer role. He was supporting an ETL application. In four years, all he had done was restart the application in case of any issues and it somehow worked for him. I've seen my fair share of such cases.
- wisecoder 6y ago99% of the companies don't pay for On call rotation / Production Support. Exploiting H1Bs for On call support is very common practice in IT industry.
- redis_mlc 6y agoCompetent DBAs have previously encountered these items for over 25 years: > because one of our database vm restarted Using VMs for production databases is always a problem because of correlated software updates and tail latency. DBAs don't like to share. Just say no. > we forgot to change the data type of one column from integer to big int(integer overflow) It's a good idea to dump your schema and grep for the key columns periodically. Then fix and monitor them. (In MySQL, there's a number of ways to rapidly use up autoinc values, including master-slave replication (esp. Galera) and upsert statements.) When I go through the RCAs for busy sites and there's no DBA, usually more than half of the outages were caused by mismanaging databases, often something as simple as a missing index. Doubly so for SV unicorns. - Coinbase is one of the most egregious.
- comeonseriously 6y agoEngineers absolutely should do prod support. But... There should be layers below. It should only come to engineering when nobody else can figure it out.