5 ms·
I fundamentally disagree with the article - sorry! I used to believe it - for many years. But as systems continue to add essential (as opposed to accidental) co
by peterbell_nyc 6y ago
I fundamentally disagree with the article - sorry! I used to believe it - for many years. But as systems continue to add essential (as opposed to accidental) complexity, the only way to run production is to run production.
Why do you want to be able to run the app elsewhere? To test new features? To reduce regressions? Reasonable goals, but at the end of the day the only thing that is identical to production is production. Unless you have a real time copy of all of the data and you continually run copies of all real time production requests, it's not production. It might be the same code and a very similar infrastructure, but without the same data and load, you are going to get regressions in production. Maybe it's flaky historic data or unexpected patterns of load, but whether we like it or not, we're all already testing in production.
I'm a huge fan of unit tests, CI, and all of the other common best practices to reduce the number of bugs that are identified in production, but you also need to have the kind of tooling and processes required to minimize recovery time and to be comfortable with testing in production. Small, easily testable, quickly shipped units of work and some flavor of feature flagging so you can dark ship code, recover quickly from outages, and do things like canary roll outs to ensure the new queries don't break with the production data at scale!
- jl6 6y agoThe tech giants are each their own unique thing and not generally a pattern that anyone else can or should follow. Most organizations, for example, do not have hyperscale volumes and face no insurmountable barriers in setting up parallel environments for testing.
- Sebb767 6y agoI disagree. Yes, fully replicating the production environment is not possible for big apps or apps with customer data, no discussion there. But when you want to debug components, test error conditions etc, a copy is extremely important. You can do "[DEBUG] added some logging" commits left, right and center, but it's not going to replace a debugger. Additionally, you might want to create custom or corrupt datasets. Yes, you can theoretically add a flag, but then this flag needs to be checked everywhere and might come with its own bugs. Using the "customers_debug" table instead of "customers" table (for example) works as well, but then you replicated staging in your production environment - added complexity which is definitely _not_ needed. Lastly, this misses the other points the article makes - a local copy allows you to freely play with the configuration, shut down related services etc. You can not do that in production. But I'll give you that - resiliency in production is still very necessary and you usually won't be able to replicate everything locally. The ability to run local isn't everything - but that was not the point the article was making.
- breatheoften 6y agoDebugging code running in the production environment with access to production resources and production versions of external services is extremely useful and not always impossible to arrange -- I think it's a reasonable expectation for good development teams to have this capability in many applications ... I was unable to run the codebase for a project I joined once and switched to using heroku one-off dynos (servers from that app's cloud provider -- these nodes are identical to production nodes excepting that they don't receive external requests and can be segregated in logs) as my development environment ... did some dumb things with ssh and tmux and was able to persuade all my tools that all the remote node's processes were accessible from localhost -- with a little finagling being able to use all the tools normally used for local development and debugging directly against production provided an extremely productive set of lenses for peering into the behavior of production resources -- especially when onboarding into the new company's codebase ...
- abgtesting 6y agoWhen your production environment involves thousands of servers and massive databases, at least one testing stage is essential. You need CI and regression tests running on smaller stages with mock data, because a bad rollout can take some time to revert, even if it didn't include a rogue database query or migration that impacted your data. Knowing how to rollback quickly is important, but a penny of prevention is worth a pound of cure with large services. And IMO you shouldn't let people "test in production" when the changes involve a production database; those should be off-limits. It's all fun and games until someone forgets which instance they're logged into. You'll still get plenty of regressions, but the goal is to avoid major issues that impact most of your users/customers/etc.
- throwdbaaway 6y agoOf course the development environment will never be the same as the production environment, due to the difference in load of the stateful services like databases and queues. I think we are all in agreement about this. For me, the article mostly concerns about the stateless services. Let me rephrase the article's main "ethos" in a different way, if you may.. If all your stateless services are running in the same kubernetes cluster in production, and you accidentally destroy the cluster (ahem terraform ahem), would you able to automatically re-deploy them from scratch? Does anyone really understand the dependencies? Is there any circular dependency which would require some hack? And if you can't do that in the local development environment, what makes you think that you can recover from a production incident like this?
- tablespoon 6y ago> Why do you want to be able to run the app elsewhere? To test new features? To reduce regressions? Reasonable goals, but at the end of the day the only thing that is identical to production is production. Unless you have a real time copy of all of the data and you continually run copies of all real time production requests, it's not production. It might be the same code and a very similar infrastructure, but without the same data and load, you are going to get regressions in production. Maybe it's flaky historic data or unexpected patterns of load, but whether we like it or not, we're all already testing in production. That is letting the perfect be the enemy of the good. Sure, there are classes of bugs you'll only be able to find in production, but there are also large classes of bugs you can find in a development/test environment. In general, it's a good thing to minimize the bugs you find in production. It's certainly less stressful.
- theknocker 6y agoYou know what really helps "minimize recovery time?" The ability to simply re-instantiate the system.
- forgotmypw17 6y agoIMHO, if you can't rebuild all of production from scratch, minus the sensitive data, the you don't have a full understanding of your production system, and that's where serious problems begin.