8 ms·
Exactly. “Staging never matches Prod” - well why is that? Make it so!!
by repler 5y ago
Exactly. “Staging never matches Prod” - well why is that? Make it so!!
- drewcoo 5y agoI have never ever even heard of a place where that was possible. The easiest way to make that scenario happen is take do whatever testing you'd have done in staging and do it in prod. Problem solved.
- saq7 5y agoI am curios, why do you think it's impossible? I think we can establish that the database is the biggest culprit in making this difficult. As an independent developer, I have seen several teams that either back sync the prod db into the staging db OR capture known edge cases through diligent use of fixtures. I am not trying to counter your point necessarily, but just trying to understand your POV. Very possible that, in my limited experience, I haven't come across all the problems around this domain.
- lamontcg 5y agoThe variety of requests and load in prod never matches production along with all the messiness and jitter you get from requests coming from across the planet and not just from your own LAN. And you'll probably never build it out to the same scale as production and have half your capex dedicated to it, so you'll miss issues which depend on your own internal scaling factors. There's a certain amount of "best practices" effort you can go through in order to make your preprod environments sufficiently prod like but scaled down, with real data in their databases, running all the correct services, you can have a load testing environment where you hit one front end with a replay of real load taking from prod logs to look for perf regressions, etc. But ultimately time is better spent using feature flags and one box tests in prod rather than going down the rabbit hole of trying to simulate packet-level network failures in your preprod environment to try to make it look as prodlike as possible (although if you're writing your own distributed database you should probably be doing that kind of fault injection, but then you probably work somewhere FAANG scale, or you've made a potentially fatal NIH/DIY mistake).
- sharken 5y agoAs if this wasn't enough of a headache, GDPR regulation requires more safeguards before you can put your prod-data in a secured staging environment. Then there is the database size, which can make it hard and expensive to keep preprod up to date. And should you want to measure performance, then no one else can use preprod while that is going on.
- nickelpro 5y agoThe article doesn't talk about any of that though. The article says staging diffs prod because of: > different hardware, configurations, and software versions The hardware might be hard or expensive to get an exact match for in staging (but also, your stack shouldn't be hyper fragile to hardware changes). The latter two are totally solvable problems
- Gigachad 5y agoWith modern cloud computing and containerization, it feels like it has never been easier to get this right. Start up exactly the same container/config you use for production on the same cloud service. It should run acceptably similar to the real thing. Real problem is the lack of users/usage.
- relaxing 5y agoA lot of things seem like they shouldn’t be, until you’ve debugged a weird kernel bug or driver issue that causes the kind of one-off flakiness that becomes a huge issue at scale.
- lamontcg 5y agoI was responding to other commentors not really the title article. The stuff you cite there is pretty simple to deal with, configuration management is basically a solved problem and IDK how you can't just fix the different hardware. The more universal problem of making preprod look just like prod so that you have 100% confidence in a rollout without any of the testing-in-prod patterns (feature flags, smoke tests, odd/even rollouts, etc) is not very solvable though.
- 5y ago
- karmasimida 5y agoIt is possible. But you need infrastructure and paying delicate attention to this problem. It is hard to define exactly what does replicating prod mean. And sometimes it might be difficult, e.g. prod might have access controlled customer data store that has its own problem, or it is about cost. But doesn't necessarily mean if you can't replicate perfectly, it is useless, you can still catch problems with things that you can replicate and do go wrong. Ofc it is impossible to catch bugs 100% with staging, however, that argument goes either way.
- anothernewdude 5y ago> I have never ever even heard of a place where that was possible. You set the CI/CD pipeline to enforce that deploys happen to staging, and then happen to production. That's it. It's not hard.
- maxbond 5y agoTech debt can definitely accrue that makes it difficult. For instance recently I had an odd difference between prod and staging, and when I looked into it I realized there was a legacy behavior for old customers, and our testing users in different environments were on different sides of the cutoff. Naturally that's a nasty tech debt and we should find a way to clean it up, but it was a pragmatic solution to a business problem that occurred before any of my team mates started working there, and it meant we could continue to serve our customers as we made a transition. These things happen.
- capn_duck 5y agoFound the guy thats only worked at a company of 1000 people
- pclmulqdq 5y agoI worked at a financial company that had an exact copy of its largest datacenter deployment in a lab for full-system testing. They recorded packets from the exchange at nanosecond precision, and had the equipment to do exact playback of those packets in the lab. It was an amazing engineering tool. You could test all sorts of end-to-end performance tweaks there, and have great confidence that you were right. I also worked at a financial company that did not have one of these. The computers ran substantially slower, and lots of end-to-end performance improvements were left on the table and mired in debate. The mock datacenter cost about $10 million, and was worth every penny.
- ehnto 5y agoI've worked at a number of places where we had multiple products with replicated staging/production environments. One approach is to have your environments codified, prod gets built from the same pipeline staging does, and to automate database refreshes from prod > staging so they don't fall behind. It isn't rocket science, but of course it doesn't come at zero cost. Some production environments are pretty hard to replicate too, like anything with third party integrations that don't offer a staging environment of their own. If you're a small shop I can understand it, but bigger companies with infrastructure teams, there's no excuse really, the technologies are all there.
- deleted 5y ago[deleted]
- throwaway6532 5y agoIf that's how it is at every single company then saying "just make it the same" probably isn't the answer.
- throwaway787544 5y agoFor many, many cases, it is literally impossible, if not incredibly unrealistic. Doing away with it eliminates all of the problems associated with it and allows you to gain more confidence in your deploys. It simplifies while adding reliability and velocity. This is why abandoning staging is the best strategy.