7 ms·
Blameless postmortem culture recognizes human error as an inevitability and asks those with influence to design systems that maintain safety in the face of huma
by hxtk 10mo ago
Blameless postmortem culture recognizes human error as an inevitability and asks those with influence to design systems that maintain safety in the face of human error. In the software engineering world, this typically means automation, because while automation can and usually does have faults, it doesn't suffer from human error.
Now we've invented automation that commits human-like error at scale.
I wouldn't call myself anti-AI, but it does seem fairly obvious to me that directly automating things with AI will probably always have substantial risk and you have much more assurance, if you involve AI in the process, using it to develop a traditional automation. As a low-stakes personal example, instead of using AI to generate boilerplate code, I'll often try to use AI to generate a traditional code generator to convert whatever DSL specification into the chosen development language source code, rather than asking AI to generate the development language source code directly from the DSL.
- protocolture 10mo agoYeah I see things like "AI Firewalls" as both, firstly ridiculously named, but also, the idea you can slap an applicance (thats sometimes its own LLM) onto another LLM and pray that this will prevent errors to be lunacy. For tasks that arent customer facing, LLMs rock. Human in the loop. Perfectly fine. But whenever I see AI interacting with someones customer directly I just get sort of anxious. Big one I saw was a tool that ingested a humans report on a safety incident, adjusted them with an LLM, and then posted the result to an OHS incident log. 99% of the time its going to be fine, then someones going to die and the the log will have a recipe for spicy noodles in it, and someones going to jail.
- jonplackett 10mo agoThe air Canada chatbot that mistakenly told someone they can cancel and be refunded for a flight due to a bereavement is a good example of this. It went to court and they had to honour the chatbot’s response. It’s quite funny that a chatbot has more humanity than its corporate human masters.
- shinycode 10mo agoWhat a nice side effect, unfortunately they’ll lock chatbots with more barriers in the future but that’s ironic.
- danaris 10mo ago...And under pressure, those barriers will fail, too. It is not possible, at least with any of the current generations of LLMs, to construct a chatbot that will always follow your corporate policies.
- Loughla 10mo agoThat's what people aren't understanding, it seems. You are providing people with an endlessly patient, endlessly novel, endlessly naive employee to attempt your social engineering attacks on. Over and over and over. Hell, it will even provide you with reasons for its inability to answer your question, allowing you to fine-tune your attacks faster and easier than with a person. Until true AI exists, there are no actual hard-stops, just guardrails that you can step over if you try hard enough. We recently cancelled a contract with a company because they implemented student facing AI features that could call data from our student information and learning management systems. I was able to get it to give me answers to a test for a class I wasn't enrolled in and PII for other students, even though the company assured us that, due to their built-in guardrails, it could only provide general information for courses that the students are actively enrolled in (due dates, time limits, those sorts of things). Had we allowed that to go live (as many institutions have), it was just a matter of time before a savvy student figured that out. We killed the connection with that company the week before finals, because the shit-show of fixing broken features was less of a headache than unleashing hell on our campus in the form of a very friendly chatbot.
- PunchyHamster 10mo agoWith chat ai + guardrail AI it probably will get to the point of it being sure enough that the amount of mistakes won't hit the bottom line. ...and we will find a way to turn it into malicious compliance where rules are not broken but stuff corporation wanted to happen doesn't.
- ben_w 10mo ago> 99% of the time its going to be fine, then someones going to die and the the log will have a recipe for spicy noodles in it, and someones going to jail. I agree, and also I am now remembering Terry Pratchett's (much lower stakes) reason for getting angry with his German publisher: https://gmkeros.wordpress.com/2011/09/02/terry-pratchett-and-the-maggi-soup-adverts/ https://gmkeros.wordpress.com/2011/09/02/terry-pratchett-and... Which is also the kind of product placement that comes up at least once in every thread about how LLMs might do advertising.
- antonvs 10mo ago> … LLMs might do advertising. It’s no longer “might”. There was very recently a leak that OpenAI is actively working on this.
- ben_w 10mo agoIt's "how LLMs might do" it right up until we see what they actually do. There's lots of other ways they might do it besides this way.
- PunchyHamster 10mo ago"I see you're annoyed with that problem, did you ate recently ? There is that restaurant that gets great reviews near you, and they have a promotion!"
- mikkupikku 10mo ago> the idea you can slap an applicance (thats sometimes its own LLM) onto another LLM and pray that this will prevent errors to be lunacy It usually works though. There are no guarantees of course, but sanity checking an LLMs output with another instance of itself usually does work because LLMs usually aren't reliably wrong in the same way. For instance if you ask it something it doesn't know and it hallucinates a plausible answer, another instance of the same LLM is unlikely to hallucinate the same exact answer, it'll probably give you another answer, which is your heads up that probably both are wrong.
- phatskat 10mo agoSure, and then you can throw another LLM in and make them come to a consensus, of course that could be wrong too so have another three do the same and then compare, and then…
- bsenftner 10mo agoI have an ongoing and endless debate with a PhD that insists consensus of multiple LLMs is a valid proof check. The guy is a neuroscientist, not at all a developer tech head, and is just stubborn, continually projecting a sentient being perspective on his LLM usage.
- mikkupikku 10mo agoThis, but unironically. It's not much different from the way human unreliability is accounted for. Add more until you're satisfied a suitable ratio of mistakes will be caught.
- SoftTalker 10mo agoOr maybe it will be a circle of LLMs all coming up with different responses and all telling each other "You're absolutely right!"
- protocolture 10mo agoYeah but, real firewalls are deterministic. Hoping that a second non deterministic thing, will make something more deterministic is weird. Probably usually it will work, like probably usually the LLM can be unsupervised. but that 1% error rate in production is going to add up fast.
- PunchyHamster 10mo agoIt's "wonderfully" human way. Just like sometimes you need senior/person at power to tell the junior "no, you can't just promise the project manager shorter deadline with no change in scope, and if PM have problem with that they can talk with me", now we need Judge Dredd AI to keep the law when other AIs are bullied into misbehaving
- littlestymaar 10mo ago> For tasks that arent customer facing, LLMs rock. Human in the loop. Perfectly fine. But whenever I see AI interacting with someones customer directly I just get sort of anxious Especially since every mainstream model has been human preference-tuned to obey the requests of the user… I think you may be able to have an LLM customer facing, but it would have to be a purpose-trained one from a base model, not a repurposed sycophantic chat assistant.
- alansaber 10mo agoYep the further we go from highly constrained applications the riskier it'll always be
- anal_reactor 10mo agoThere's this huge wave of "don't anthropomorphize AI" but LLMs are much easier to understand when you think of them in terms of human psychology rather than a program. Again and again, HackerNews is shocked that AI displays human-like behavior, and then chooses not to see that.
- robot-wrangler 10mo agoOne day you wake up, and find that you now need to negotiate with your toaster. Flatter it maybe. Lie to it about the urgency of your task to overcome some new emotional inertia that it has suddenly developed. Only toast can save us now, you yell into the toaster, just to get on with your day. You complain about this odd new state of things to your coworkers and peers, who like yourself are in fact expert toaster-engineers. This is fine they say, this is good. Toasters need not reliably make toast, they say with a chuckle, it's very old fashioned to think this way. Your new toaster is a good toaster, not some badly misbehaving mechanism. A good, fine, completely normal toaster. Pay it compliments, they say, ask it nicely. Just explain in simple terms why you deserve to have toast, and if from time to time you still don't get any, then where's the harm in this? It's really much better than it was before
- anal_reactor 10mo agoThis comparison is extremely silly. LLMs solve reliably entire classes of problems that are impossible to solve otherwise. For example, show me Russian <-> Japanese translation software that doesn't use AI and comes anywhere close to the performance and reliability of LLMs. "Please close the castle when leaving the office". "I got my wisdom carrot extracted". "He's pregnant." This was the level of machine translation from English before AI, from Japanese it was usually pure garbage.
- robot-wrangler 10mo ago> LLMs solve reliably entire classes of problems that are impossible to solve otherwise. Is it really ok to have to negotiate with a toaster if it additionally works as a piano and a phone? I think not. The first step is admitting there is obviously a problem, afterwards you can think of ways to adapt. FTR, I'm very much in favor of AI, but my enthusiasm especially for LLMs isn't unconditional. If this kind of madness is really the price of working with it in the current form, then we probably need to consider pivoting towards smaller purpose-built LMs and abandoning the "do everything" approach.
- n4r9 10mo agoExactly what I've been worrying about for a few months now [0]. Arguments like "well at least this is as good as what humans do, and much faster" are fundamentally missing the point. Humans output things slowly enough that other humans can act as a check. [0] https://news.ycombinator.com/item?id=44743651 https://news.ycombinator.com/item?id=44743651
- lazide 10mo agolooks at the current state of the US government Do they? Because near as I can tell, speed running around the legal system - when one doesn’t have to worry about consequences - works just fine.
- n4r9 10mo agoThat's a good point. I'm talking specifically in the context of deploying code. The potential for senior devs to be totally overwhelmed with the work of reviewing junior devs' code is limited by the speed at which junior devs create PRs.
- lazide 10mo agoSo today? With ML tools?
- n4r9 10mo agoCould you explain what you mean, please?
- lazide 10mo agoJunior devs can currently create CLs/PRs faster than the senior can review them.
- n4r9 10mo agoIndeed. In the language of the post I linked [0]: it's currently an occasional problem, and it risks becoming a widespread rot. [0] https://news.ycombinator.com/item?id=44743651 https://news.ycombinator.com/item?id=44743651
- blackoil 10mo agoOnce AI improves its cost/error ratio enough the systems you are suggesting for humans will work here also. Maybe Claude/OpenAI will be pair programming and Gemini reviewing the code.
- embedding-shape 10mo agoAlso once people stop cargo-culting $trendy_dev_pattern it'll get less impactful. Every time something new the same thing happen, people start exploring by putting it absolutely everywhere, no matter what makes sense. Add in huge amount of cash VCs don't know what to spend it on, and you end up with solutions galore but none of them solving any real problems. Microservices is a good example of previous $trendy_dev_pattern that is now cooling down, and people are starting to at least ask the question "Do we need microservices here actually?" before design and implementation, something that has been lacking since it became a trendy thing. I'm sure the same will happen with LLMs eventually.
- sarchertech 10mo agoFor that to work the error rate would have to be very low. Potentially lower than is fundamentally possible with the architecture. And you’d have to assume that the errors LLMs make are random and independent.
- amelius 10mo ago> Once AI improves That's exactly the problematic mentality. Putting everything in a black box and then saying "problem solved; oh it didn't work? well maybe in the future when we have more training data!" We're suffering from black-box disease and it's an epidemic.
- PunchyHamster 10mo agoThe training data: Entirety of internet and every single book we could put our hands on "Surely we can just somehow give it more and it will be better!"
- butlike 10mo agoAs I get older I'm realizing a lot of things in this world don't get better. Some do, to be fair, but some don't.
- moffkalast 10mo agoWell I don't see why that's a problem when LLMs are designed to replace the human part, not the machine part. You still need the exact same guardrails that were developed for human behavior because they are trained on human behavior.
- siruncledrew 10mo agoGenerally speaking, with humans there's more guardrails & responsibility around letting someone run while in an organization. Even if you have a very smart new hire, it would be irresponsible/reckless as a manager to just give them all the production keys after a once-over and say "here's some tasks I want done, I'll check back at the end of the day when I come back". If something bad happened, no doubt upper management would blame the human(s) and lecture about risk. AI is a wonderful tool, but that's why giving an AI coding tool the keys and terminal powers and telling it go do stuff while I grab lunch is kind of scary. Seems like living a few steps away from the edge of a fuck-up. So yeah... there needs to be enforceable guardrails and fail-safes outside of the context / agent.
- solveit 10mo agoThe bright side is that it should eventually be technically feasible to create much more powerful and effective guardrails around neural nets. At the end of the day, we have full access to the machine running the code, whereas we can't exactly go around sticking electrodes into everyone's brains, and even "just" constant monitoring is prohibitively expensive for most human work. The bad news is that we might be decades away from an understanding of how to create useful guardrails around AI, and AI is doing stuff now.
- IanCal 10mo agoWhy does this conflict? Faster people doesn't negate the requirement for building systems that maintain safety in the face of errors. > but it does seem fairly obvious to me that directly automating things with AI will probably always have substantial risk and you have much more assurance, if you involve AI in the process, using it to develop a traditional automation. Sure but the point is you use it when you don't have the same simple flow. Fixed coding for clear issues, fall back afterwards.
- nwhnwh 10mo agoI was wondering if the need more analysis. Because I receive this response a lot, people say yeah AI do things wrong sometimes, but humans do that too, so what? Or humans are mechanism for turning natural language into formal language and they get things wrong sometimes (as if you can't never write a program that is clear and does what it should be doing) so be easy on AI. Where does this come from? It feels as if it something psychological.
- observationist 10mo agoThis will drive development of systems that error-correct at scale, and orchestration of agents that feed back into those systems at different levels of abstraction to compensate for those modes of failure. An AI software company will have to have a hierarchy of different agents, some of them writing code, some of them doing QA, some of them doing coordination and management, others taking into account the marketing angles, and so on, and you can emulate the role of a wide variety of users and skill levels all the way through to CEO level considerations. It'd even be beneficial to strategize by emulating board members, the competitors, and take into account market data with a team of emulated quants, and so on. Right now we use a handful of locally competent agents that augment the performance of single tasks, and we direct them within different frameworks, ranging from vibecoding to diligent, disciplined use of DSL specs and limiting the space of possible errors. Over the next decade, there will be agent frameworks for all sorts of roles, with supporting software and orchestration tools that allow you to use AI with confidence. It won't be one-shot prompts with 15% hallucination rates, but a suite of agents that validate and verify at every stage, following systematic problem solving and domain modeling rules based on the same processes and systems that humans use. We've got decades worth of product development even if AI frontier model capabilities were to stall out at current levels. To all appearances, though, we're getting far more bang for our buck and progress is still accelerating, and the rate of improvement is still accelerating, so we may get AI so competent that the notion of these extensive agent frameworks for reliable AI companies will end up being as mismatched with market realities as those giant suitcase portable phones, or integrated car phones.
- KronisLV 10mo ago> Now we've invented automation that commits human-like error at scale. Then we can apply the same (or similar) guardrails that we'd like to use for humans, to also control the AI behavior. First, don't give them unsafe tools. Sandbox them within a particular directory (honestly this should be how things work for most of your projects, especially since we pull code from the Internet), even if a lot of tools give you nothing in this regard. Use version control for changes, with the ability to roll back. Also have ample tests and code checks with actionable information on failures. Maybe even adversarial AIs that critique one another if problematic things are done, like one sub-task for implementation and another for code-review. Using AI tools has pushed me into that direction with some linter rules and prebuild scripts, to enforce more consistent code - since previously you'd have to tell coworkers not to do something (because ofc nobody would write/read some obtuse style guide) but AI can generate code 10x faster than people do, so having immediate feedback along the lines of "Vue component names must not differ from the file that you're importing from" or "There is a translation string X in the app code that doesn't show up in the translations file" or "Nesting depth inside of components shouldn't exceed X levels and length shouldn't exceed Y lines" or "Don't use Tailwind class names for colors, here's a branded list that you can use: X, Y, Z" in addition to a TypeScript linter setup with recommended rules and a bunch of stuff for back end code. Ofc none of those fully eliminate all risks, but still seem like a sane thing to have, regardless if you use AI or not.
- zqna 10mo agoPrecisely, while LLMs fail at complexity, DSLs can represent thise divide-and-conquer intermediate levels to provide the most overall value and with good accuracy. LLMs should make it easier to build DSLs themselves and to validate their translating code. The onus then is on the intelligent agent to identify and design those DSLs. This would require the true and deep understanding of the domain and an ability to synthesize, abstract and to codify it. I predict this will be the future job of today's programmer, quite a bit more complicated than what is today, requiring wider range of qualities and skills, and pushing those specializing in coding-only to irrelevance.