4 ms·
That's a good idea, but a physical device is deterministic most of the time (if not always). E.g.: A lawnmower, as credited by the great Bryan Cantrill. Howeve
by bayindirh 18d ago
That's a good idea, but a physical device is deterministic most of the time (if not always). E.g.: A lawnmower, as credited by the great Bryan Cantrill.
However an AI agent, or the model powering it is stochastic by design. How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%?
BTW, really, how is that AI observability work is going in the frontier labs? Do they care, even?
- ragebol 18d agoThat we don;'t understand it is not an excuse, it's all the more reason to not let these things roam freely, with this amount of potential to do damage.
- bayindirh 18d agoWe're in agreement, then. :)
- tessierashpool 18d agoHow can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%? That's a question any lawmaker has already had to ask about technology all the time. I'm not saying they came up with great answers, but there's nothing qualitatively new about that. The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law.
- bayindirh 18d ago> The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law. We're on the same page. What I'm saying that certifying them as safe is harder than certifying a drill as safe, and we shall be more cautious about AI related technology and be more stringent about the can of worms it opens without hesitation.
- CPLX 18d agoThat's a ridiculous distinction. Is AI less deterministic than an airline dealing with weather? Of course not. The difference is one of those two things has a culture of safety and is well regulated, and the other one isn't.
- bayindirh 18d agoAn airline has a weather radar which shows the same thing for the same thing of weather event ahead. So, for similar weather phenomena, radar shows a similar thing. For that thing, procedures and regulations are built. So regulations fit into a well understood phenomena, incl. "return back because that thing is way powerful for us". For the same prompt, an AI model can return two completely different outputs, incl. but not limited to content, length, formatting and tiny details. What you get is a single instance. So, regulating an AI model for safety or any other property is not as easy as regulating air travel. Moreover, you have much stronger motivations for regulating airlines. Otherwise people die in a visible and gruesome way. With AI, it's easy to whitewash problems. Somebody committed suicide? "They were already unstable". AI told something wrong and created problems? "The tech can’t guarantee truth because it's not alive, it can't understand right and wrong". It did something good? "It's probably a sentient being, we shall respect them". I'm for regulating these things. They are dangerous as they are useful (sometimes), but the forces and motivations for regulating it is not the same.
- CPLX 18d agoOf course, it's the same. It's computer software. It's an incredibly powerful business automation tool. It's a lot of things. What it's not is God or an independently conscious entity that somehow trumps a thousand years of common law that's built up until now about torts and liability. Of course, there are some novel issues here that'll pop up here and there, but the idea that this is fundamentally different is propaganda on the part of these AI labs because the more boring, obvious situation doesn't favor them.
- saghm 17d ago> What it's not is God or an independently conscious entity that somehow trumps a thousand years of common law that's built up until now about torts and liability. Agreed, the tendency of people on tech to assume that whatever the most recent thing we've come up with is unprecedented and shouldn't have to follow all of the established patterns we've built up in society for making things safe is wild. I don't know what the next Big Thing will be but I'm pretty confident there will be people claiming it's so different from everything before that we have no choice but to throw out all of the rules for it in the name of progress.
- vipshek 18d agoI agree with one of the sibling comments that determinism isn't necessary for certifying a product. All engineered products operate under uncertain conditions; we define standards for how those products ought to respond under those conditions and verify them under measurement. Consider robot vacuums, for example. I also agree that qualitatively, this technology seems different than the others. However, I feel that people tend to overly fixate on their internal stochasticity. Even if LLMs' internal mechanism is nondeterministic, shouldn't we be able to verify their "side effects" aren't harmful? Of course, "harm" is subjective and at this scale, the most effective way to verify behavior is probably some kind of LLM-as-judge... Anyway, in this case the problems have occurred while actually running the evals themselves, so again, we're in a situation where we can't even confidently test these things and know that they won't cause harm in the outside world.
- fultonn 18d ago> How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%? By verifying that all of its possible behaviors conform with the "it works" spec, regardless of which of those behaviors it chooses. Monitoring with a known-safe fallback is the easiest case.
- intended 18d agoIf something is risky, and the end operator cannot be considered to have orchestrated the outcomes of too use, then that tool typically has significant restrictions placed on it. Liability will shift to the maker of the tool if they claim that it’s easy to use, safe, or that you don’t need unique skills or training to use it. That would be considered reckless. Cars analogy - We have licenses for cars, and different types for different vehicle classes. Cars have to be rigorously tested to meet standards to be considered road safe.
- tetha 17d agoIt's going to be interesting, because liability cases tend to revolve around the involved people, the duty they had in a situation, and if they fulfilled that duty (or were prevented in some way by someone else not fulfilling their duty). For example, for a runaway car (example from a sibling comment), the driver could be liable because they forgot the parking brake. The driver could be liable for a lack of maintenance and inspection. A mechanic could be liable for not reinstalling brake pads correctly. Or the manufacturer of the car or the brake pads could be liable because of a systemic defect. Or it could grow even more complex, maybe the brakes are designed that they have to be maintained in a very specific way, and the mechanic did a reasonable maintenance and inspection but it failed later due to this maintenance. That could split liability between the manufacturer and the mechanic. As an example, with other software, you as a developer or operator of a software have a duty to ensure it does not access computer systems you do not own in unintended ways. And this could go beyond liability into criminal territory. It'll be interesting what OpenAI gets slapped with there.
- CJefferson 18d agoI could make a non-deterministic chainsaw fairly easily. I’d also get sued into the ground if I sold it, and I wouldn’t be able to claim ‘Oh, it’s just an unavoidable part of progress’.
- nextaccountic 18d agoA non-deterministic machine is usually called defective
- alwa 17d agoIt’s not a defect, the stochastic “temperature” setting unleashes the chainsaw’s creativity and imagination! Who are we to cast aspersions on the Oracle Chainsaw’s intelligence—nay, wisdom!—just because it happens to be non-living?
- Kim_Bruning 17d agoA chainsaw is a physical machine. Physical machines are technically non-deterministic if you look closely. They have Variance. The discipline to manage variance is called Tolerance. Many physical machines and components come with a datasheet that will list their tolerances. Failure to correctly document tolerances does in fact get you sued. However, while this is truly a great idea, we're not going to be able to make it work for computational systems. Computers, software, and also LLMs are sensitive to initial conditions. Which is why tolerances are not so familiar to computer people. (but not entirely: eg your PSU might list 110-240Vac/300W as input tolerance) Interestingly, LLMs actually have a somewhat lower sensitivity to initial conditions than traditional interpreters. See what happens if you misspell "What is One Plus nOe?". So they're actually a skosh off the edge and towards the middle, though I'd argue still very much at the computational end, just from the sheer scale of the valid inputs and outputs. Mind you, if you have a pretrained LLM doing a measurable task on a line, possibly some sort of tolerances could be determined. Not so much when doing arbitrary chat. Something unintuitive: I bet that often setting the temperature > 0 (aka introduce stochasticity deliberately, variously comparable to dithering or simulated annealing in other disciplines - doing the thing where you escape local minima) will tighten the output tolerance range and improve reliability, especially in iterated processes. This works for a lot of physical and digital processes actually, and LLMs simply stole the same trick. (edit: I'm trying to compress a huge chunk of dynamics intuition in a few lines here. Hopefully still useful. TL:DR; Everything real is continuous and noisy if you look close; and you're really trying to build attractors and bound variance, if you can. )
- Melatonic 18d agoSounds like owning a Dog
- CPLX 17d agoExactly people have owned and deployed animals in the world for millennia and actually a huge chunk of common law was developed precisely to deal with the various unpredictable events that resulted and harms caused to others. Like has anyone heard of horses? The idea that AI can’t possibly be addressed because it could autonomously break free and ruin something is fucking ridiculous.
- saghm 17d agoHow confident are you that when you create a new UUID, it won't collide with one of your existing ones? My guess is that even though you don't get the same on every time, you're extremely confident that getting a duplicate is a extemely rare edge case that might happen in large volume but mostly isn't a concern, and furthermore, I'm guessing you understand that the risk can still be quantified. Casinos can't make slot machines that literally never pay out, but it's a different result every time you pull the lever. We have existing legal frameworks for how to regulate things that aren't perfectly predictable (an economist might argue that if it were possible to predict slot machines then casinos with them would all go out of business).