6 ms·
What seasoned companies do for reducing bugs in production
Long story short: me and my team have been forced to develop a software in an extreme hurry.
This happened because this is a software that produces money for the company and the top management has a poor vision about preventing problems.
The result is a really bad codebase, difficult to maintain and especially difficult to test.
At the same time this codebase is handling critical and dangerous stuff (like bills and monthly charges), and we have already thousands of customers.
The software has been launched 3 months ago.
Since I'm responsible for this software (I got promoted to manager), I want to implement some serious processes for reducing, as much as I can, the chances to introduce bugs.
As now we've introduced some peers code review, automatic testing on most critical stuff (but since the codebase sucks these aren't really reliable tests) and extensive manual tests before every release. We also become better at fighting project managers and other managers on their continue pressures.
I look for some advice on what I can do for reducing the possibility of introducing some serious bugs in the code.
What are the things that seasoned software companies do?
I'll do my best for convincing the company to slow down new features and releases for implementing safer processes.
The management has a "move fast" and high-pressure culture so doing this will be hard but that's another story, I'll do my best.
Thank you for your advise
- java-man 3y agoIn my experience, in the order of importance: 1. hire qualified people 2. document all the requirements and business decisions 3. rigorous design reviews 4. mandatory code reviews 5. implement tests for each requirement 6. have a regression testing suite 7. have a user acceptance test environment, separate from production
- odle 3y agothank you, I'll try to implement each point
- mattbrewsbytes 3y agoDo you have test engineers or people that write automated tests? Not unit tests or tests alongside the code but tests that test say a development environment and are run outside the code base. Building an automated test suite like that as you build the systems is key to making sure past stuff built still works with new stuff. There are also tools you can use to analyze the code and provide suggestions for refactoring, if devs use this as they build things then it should help along the way. Sneak fixing tech debt in as you go. Get a CI/CD workflow going to run tests on branches, after deploys, etc. and keep building the test suite out. When you have larger refactor work identified try to get a really tight and well defined scope - PM’s have fears of engineers just wanting to go off and perfect things, rightly so because it happens.
- deleted 3y ago[deleted]
- odle 3y agoAs now we've only integration and unit test within the codebase but I'll try to develop also external functional test, it would be great. About code analysis tools, I'll definitively check them out! I understand the PM fears, but I clealy pointed out to PMs that we worked like monkeys (with regular massive overtime) for shipping this product. Just being allowed to provide time estimates when the next bussiness critical feature comes would be great
- softwaredoug 3y agoI prefer a "quarantine" approach testing giant ugly legacy code bases. You isolate as much of the code as possible behind a well-understood and relatively stable interface. Sometimes that's an API boundary unlikely to change, where you standup a test instance of the app and database. Then begin testing there. Standup a version of the app, test it end-to-end, require all tests to pass before releases. When you fix a bug, ensure a test in this suite comes with it. Then with that secure perimeter in place, you can (if you so choose / need to) move in to the buggiest parts of the code to refactor and introduce better unit testing.
- odle 3y agoyeah, I found that to be a good approach, we're doing the same thing where possible. Thank you
- Jemaclus 3y agoMy advice as a manager is as follows: 1) You always have time for tests. Your estimates should include tests and testing. If someone asks for a feature and you think it would take 3 days to test, half a day to write tests, and half a day to do manual testing, then it will take 4 days. That's your estimate. You always have time for tests. Tell your team this with every ticket. Don't sandbag the estimates, but make sure you are planning time for your team to ensure high quality. 2) One of your many jobs as a manager is to protect your team from external pressure. Think of yourself as the submarine (too soon?) that keeps your team safe from the pressures of the business. The team should be focused on shipping high quality features, not whether the CEO is asking for something or not. If a feature needs more time for testing, it's your job to take the heat for it. Both of these will give your team the breathing space to actually write tests and actually test their code, without fear of reprisal and without fear of losing their jobs. Right now, they probably have very little trust in this process. It's your job (another one!) to prove to them that they can trust you to protect them as they do their jobs. Over time, the team will develop a culture of testing and releasing higher quality products. You must remain consistent. You should hold the line, hold your team to a higher standard than exists now, and hold your external stakeholders at bay. Additionally, I personally don't believe in deadlines. With few exceptions such as legal/court orders or certain contractual obligations, deadlines are arbitrary dates made up by someone for some arbitrary reason. From a financial or company survival standpoint, there's very little difference in real terms between releasing a feature on Tuesday and on Wednesday or Thursday or even Friday. I try to think of these as target dates. We will do our best to meet them, but if we need an extra day to test, then we'll take it. Hope that helps. Good luck.
- odle 3y ago1) and 2) I can't agree more. However, your guess is right, the team have little trust in the process. Now that there aren't really business critical new features, I think I'm successully not giving them pressure. But, as already happened in other teams/context, when the time for an important business feature comes, the director/business could "override" me and giving the team extreme pressure. About the deadline, I agree with you as a professional, but in this company they like to organize launch keynotes and stuff like that so we need to ship within the keynote date. (vent mode) You wonder how the keynote date is set? I don't really know, with my boss I really don't see any transparency in that
- Jugurtha 3y ago>As now we've introduced some peers code review, automatic testing on most critical stuff (but since the codebase sucks these aren't really reliable tests) They may not be "reliable", but these are your safety net, or harness, so you don't fall. I wrote about similar issues, for instance here: https://news.ycombinator.com/item?id=26591067 https://news.ycombinator.com/item?id=26591067 and, given your promotion, here: https://news.ycombinator.com/item?id=37211796 https://news.ycombinator.com/item?id=37211796. It contains a few steps starting from "So...". You can add monitoring, something like Sentry (https://sentry.io https://sentry.io) will capture exceptions that were not handled that you have not seen because the stack trace is buried in hundreds of pages of logs or something. It groups them by exception and counts them. It's pretty awesome. (https://docs.sentry.io https://docs.sentry.io). It supports around 108 platforms (Java, Python, JavaScript, etc.). This lets you see the exceptions and makes prioritizing easier (which ones are the most frequent, which ones impact the most, etc.). If you don't have them already, issue templates are really useful and the comment I linked to explains why, but here's an example of an issue template (again, you can configure them for different types of issues so team members select from a dropdown for a bug or a feature): <!--These are comments and will not be rendered. Please click Preview to check.--> ### Description ---------------- <!--Describe the problem we'll solve. What, where, when, who, why, and how.--> ### Target ---------- <!--Who's the target audience? Who'll enjoy this the most or is suffering from this bug or lacking this feature the most? Is this a problem you have? Have you talked with other people who have this problem? Maybe it's something that would just be cool? Is it a seed for an idea that could be cool? --> ### Instances ------------- <!--Can you describe the times this problem has happened for you or others--> ### Proposal ------------ <!--What's the plan to solve this problem? Do we have a candidate solution or a hunch? A list of steps? It's okay if they are a bit vague for now. We'll clarify by debating them and adding context to them, reading more about the subject. It's essential to understand that this document serves to capture thoughts and externalize thought processes so we can all tackle the problem more effectively. --> ### Resources ------------- <!--Any recommended reading about this? Videos to watch about something similar? Maybe a blog post, documentation pages, a code snippet? Sketch, images, mock-ups, etc... --> /label ~feature /cc @odle <!-- You will get an email with every event on that issue -->
- muzani 3y agoYou can also work on damage control. It's ideal to tell everyone to test thoroughly and never release bugs. But if you're stuck in a high bug environment, it can be useful to set up a system where you can fast track bug fixes to production in minutes, sometimes even before the customer realizes you were down. Especially in a startup environment, you can take advantage of the high speed, low quality culture to just release fixes faster. Sometimes a lot of damage comes in when a fix takes too long to approve - billing being down for hours, for example. We have people on a rotated shift every week, aka fire wardens as well as a manager on emergency duty. There should be channels to contact them for emergencies, personal numbers, etc. The fire warden's job isn't necessarily to fix bugs but rather to direct them to the best person. The release engineer is also the fire warden, to prevent being blocked by each other. There should also be a clear process of to release emergency fixes, e.g. CS talks to manager. Manager talks to engineer. Engineer releases fix. Engineer pushes fix to staging. QA runs the bare minimum of testing. Engineer releases it to prod. If you can't do that quickly, find out what's blocking it. Figure out what the bare minimum of testing is. Maybe consider CI/CD if it works for you. We've realized that Git Flow got in the way too and used trunk based development helped a lot in releasing faster: https://trunkbaseddevelopment.com/ https://trunkbaseddevelopment.com/
- sloaken 3y agoYears ago I worked for a company that made critical equipment. We did 3 things: 1) planned tests with test scripts and dedicated Test engineers. 2) Software intensive inspection - team of 4 or 5. A) Team lead controls the process. B) Note taker - this is the author of the code C) Reader - reads the code and says what they think it is doing D&E) supplemental people who sometime ask questions. It is a VERY slow process. Person B has to be sure not to take it personally, as they feel assaulted. Person A helps person B not feel bad. Meeting should NEVER last more than a hour, as it is too hard on brains and emotions. This catches a LOT of things. 3) Regression testing (automated) with the goal of code coverage. Start by testing all success paths and then fail paths. If you get 60% coverage I believe that is a HUGE amount. Run the tests daily? or when chages are proposed.
- odle 3y agoWhen I'll have the "power" to do that, I'll ask for dedicated tests engneers. As now, we can implement some scripts (external to the codebase testing suite) for e2e testing. About the code inspection and regression testing, I'll try to apply what you suggested, thank you.
- sloaken 3y agoThe things that drove all three was a combination of lawsuits, and maintenance. When you start 'INHERITING' code from someone, whom you thought was only an OK developer, but has moved on to a better job, and you see the CRAP they leave behind, and you figure the amount of time you have to waste fixing their ganky code, it is easy to be bitter about the company not spending money upfront to ensure standards. Not that I am bitter :-)
- cheryledward217 3y ago[dead]