4 ms·
> But the idea that catching a bug in production is a huge problem is a myth You clearly have never worked in fintech, banks or payments. I personally witnesse
by jnardiello 10y ago
> But the idea that catching a bug in production is a huge problem is a myth
You clearly have never worked in fintech, banks or payments. I personally witnessed the moment we catched a race condition which costed the company n-k euros in missed transactions.
And this isn't even considering mission critical software where bugs can (literally) kill.
- dimino 10y agoRelax, I don't think I said bugs are never bad, or devastating. I said the idea that catching a bug in production is a huge problem is a myth. It is not, by default, true. It is true given additional information for a specific case, but is not the general rule.
- jerf 10y agoBug impacts are highly non-linearly and non-Gaussianly distributed. It doesn't matter if the "average" bug is not that big a deal to find in production, what matters is what the worst thing you can find in production is. Even if the study is not literally correct in the average case, it doesn't take much modeling or much real-world experience to see one is still wise to develop with very similar ideas in mind, if not an even more intense focus on getting bugs identified before production because the paper probably understates the impact given real bug distributions. (I say "identified" because you don't always fix identified bugs. But you really, really want to have as much as possible identified before production. Production is a terrible place to find bugs for the first time. Yeah, every once in a while an identified bug might have a far worse than expected impact, I've had that happen, but at least in my experience the completely unidentified bugs are what kill you the hardest.)
- dimino 10y agoThis is an older style of thinking that is, while safe, slower. Competition who doesn't follow this model and pushes fast, failing often will generally run faster to market than a company who follows this line of thinking. Their software might be buggier, but the concept of "good enough" applies. If you're building rocket engines that carry people, you think like this. If you're building a social networking website, or a food ordering website, or a home sharing website, the time it takes to go from "oh there's a bug" to "that bug is fixed in production" is a matter of hours, if not minutes.
- jerf 10y agoThe problem with bugs is not the bugs. It's what the bugs do. Bork up your database and the bug fix is not "minutes". (Especially if you don't notice for a while.) Piss off your customers and the bug fix is not "minutes". Screw up your money handling and the fix is not "minutes". (Well, your part of the fix may be just one minute, though that minute will be "you're fired".) If you've only ever encountered bugs that can be fixed in "minutes", you're either very lucky, or not working on anything all that important. Or possibly, not very perceptive and you've actually got a mess on your hands and you haven't realized it yet. I've inherited systems run by such people; they think everything's hunky dory but it turns out you can hardly figure out how to connect two records in the database together correctly because they were too smart for academic bullshit like referential integrity, and, lo, their database was low on the integrity. Tends to work, except the amount of elbow grease required increases without bound until it exceeds the capabilities of the developer in question, and suddenly, one day, they wake up and their job is in serious jeopardy because of what's going down and what they can't get back up. You assert that this is the "older" style of thinking, which shows a lack of understanding of your computer programming history. What you advocate is "cowboy" thinking, and it predates what I'm discussing by quite a bit. The entire 1970s was basically run this way, and it shows in those handful of remaining technologies that are still with us, and their tendency to just pound through problems as if errors are an admission of failure.
- dimino 10y agoI think you're having a reaction to the higher risk profile that comes with breaking in production. That's understandable, but wrongheaded. Plenty of companies are very successful with the mindset as I have described. Netflix runs their Simian Army in production, for example. They sometimes pull infrastructure pieces offline intentionally to test their resilience. A good production infrastructure is incredibly resilient.
- deleted 10y ago[deleted]
- JabavuAdams 10y agoSo, excluding life-critical software, it's still a business decision. Engineering will naturally tend to geek out on infrastructure projects like this and they need to be kept focused on the business case. There has to be some push-back. What is the cost of even one engineer working full-time on build verification? The thing you really want to avoid is getting blind-sided by devastating bugs or systemic process problems. So it's a balance. I've just seen many cases where projects got bogged down as developers built their super-uber-build-test framework. It's easy for people to push these projects through without rationally investigating the cost / benefit, because ... because ... you're not seriously suggesting that we shouldn't test more, you monster!?