2 ms·
A lot of the pain around production support is easily solved by having staff in multiple time zones. A good structure is to have first line support be relative
by jake_morrison 6y ago
A lot of the pain around production support is easily solved by having staff in multiple time zones.
A good structure is to have first line support be relatively generic ops people. They can handle problems related to infrastructure, e.g. hardware failures, network problems, or issues that can be handled by adding resources. The deployment process should be consistent enough across applications that they can e.g. roll back to a previous release.
This covers the majority of production problems. After that, it's time to bring in someone who understands the details of how the application works. If the dev team is geographically distributed, then someone is available during working hours. Otherwise, we have to get someone out of bed.
If the dev team has done their job right, this should be a rare occasion. Making the dev team fully responsible for the reliability of the application means that they are motivated to make it reliable. Otherwise there is a tendency to have an underclass of ops people who get abused.
A fundamental mindset here is taking responsibility for the user experience, including reliability. If this is not owned by the product development team, then who?
- mtberatwork 6y agoConvincing folks at the top of food chain that more staff is needed is one of the most difficult things to do.