5 ms·
One could only speculate on the cause here without more info. I am reminded of reading The Toyota Way a while back, which does talk about how they would shut do
by rtpg 3y ago
One could only speculate on the cause here without more info. I am reminded of reading The Toyota Way a while back, which does talk about how they would shut down lines for any defect. A coworker of mine was pretty adamant that this was A Good Idea(TM), but I thought about how much of IT best practices are about uptime and doing (often principled!) workarounds when you have operational issues, and moving forward in a degraded state.
I wonder if there are IT operations that do try to aim for the ~no defect approach like this.
- housemusicfan 3y agoToyota sounds like the kind of place that still runs Solaris on SPARC and has no immediate plans to change.
- r00fus 3y agoYou jest but back in the day (90s) one of my friends said they still ran Smalltalk VMs in production. That was 25 years ago but still possible…
- deleted 3y ago[deleted]
- kenhwang 3y agoFrom my conversations with some Toyota software engineers, they have a respectably sophisticated modern software/tech stack. There's a ton of C++ since they have to interface with hardware. For tooling and internal applications it's a lot of Rails/Elixir/Javascript. Web APIs are backed by Elixir/Go. Data processing in Scala/Spark. Infrastructure in AWS, orchestrated with k8s. Truly agile prototyping then polish, significant automated testing. Honestly, it makes "big tech" look like dinosaurs with their megasized-Java projects and 3:1 PM:engineer ratio waterfall planning once a year releases. What they're missing is the operational experience with operating the software at scale.
- upon_drumhead 3y agoI’m sure finance folks are quick to hit the big red stop button if it’s causing invalid transactions to clear. But generally the impact to the end user doesn’t really matter. There’s no reason to shut it all down if 1% of your request to a web server gets a 500. It might even make it harder to understand.
- onion2k 3y agoThere’s no reason to shut it all down if 1% of your request to a web server gets a 500. This is more like finding a bug in production and having all the devs stop working (including merging more changes) until it's fixed. Customers can still use the end product, you just stop making new things until you've solved the problem. There's some undeniable logic to doing that if you value bug free over fast.
- FirmwareBurner 3y ago>I am reminded of reading The Toyota Way a while back, which does talk about how they would shut down lines for any defect. Like a crack in the frame?[1] Shots fired! [1] https://news.ycombinator.com/item?id=37292321 https://news.ycombinator.com/item?id=37292321
- stubish 3y agoA lot of the Toyota methodologies ended up in various agile software practices. Not so much operations of a live system, but development practices. For example, all your teams push their changes to an integration branch which is automatically deployed to staging systems for testing and automatically deployed to production. And when a failure is found, that pipeline gets shut down until the problem is solved. And all eyes are looking at that problem, rather than tapping away in their own world ignoring the 'somebody else's problem'. Nobody is pushing new code unrelated to fixing the problem, making the testing and deployment of an update slower. And in theory, improves your uptime because the quality of code ending up on production is better, and when that fails, fixes will end up on production faster.
- rtpg 3y agoThat's interesting, I had never considered "main branch is red" as a similar fail state to the Toyota case, but I totally agree with that comparison!
- jacquesm 3y agoInstead you should do the opposite: aim to get the process completed even in the eye of failures. Because failures at all levels, including hardware are a fact of life in computing, more so in distributed computing and almost every problem is now a distributed problem.
- rtpg 3y agoI would like to pose the claim that distributed computing is probably not more complicated than the distributed production of automobiles. Maybe The Toyota Way is actually bunk from the get go. But if you buy into the idea that it does work, then why would it not be applicable in software development? This is the core thing that interests me.
- monomers 3y agoIn IT we (wrongly) use the word "production" to refer to the systems in operation serving customers: ie. the car that has left the factory. I don't know much about cars, but Toyota has a reputation for high reliability there. In manufacturing, production lines instead refer to a previous step in the lifecycle, still in the factory. That's where you can pull the andon cord.