4 ms·
If you can't stop an LLM from _saying_ something, are you really going to trust that you can stop it from _executing a harmful action_? This is a lower stakes p
by eximius 1y ago
If you can't stop an LLM from _saying_ something, are you really going to trust that you can stop it from _executing a harmful action_? This is a lower stakes proxy for "can we get it to do what we expect without negative outcomes we are a priori aware of".
Bikeshed the naming all you want, but it is relevant.
- nemomarx 1y agoThe way to stop it from executing an action is probably having controls on the action and an not the llm? white list what api commands it can send so nothing harmful can happen or so on.
- Scarblac 1y agoIt won't be long before people start using LLMs to write such whitelists too. And the APIs.
- omneity 1y agoThis is similar to the halting problem. You can only write an effective policy if you can predict all the side effects and their ramifications. Of course you could do like deno and other such systems and just deny internet or filesystem access outright, but then you limit the usefulness of the AI system significantly. Tricky problem to be honest.
- swatcoder 1y ago> If you can't stop an LLM from _saying_ something, are you really going to trust that you can stop it from _executing a harmful action_? You hit the nail on the head right there. That's exactly why LLM's fundamentally aren't suited for any greater unmediated access to "harmful actions" than other vulnerable tools. LLM input and output always needs to be seen as tainted at their point of integration. There's not going to be any escaping that as long as they fundamentally have a singular, mixed-content input/output channel. Internal vendor blocks reduce capabilities but don't actually solve the problem, and the first wave of them are mostly just cultural assertions of Silicon Valley norms rather than objective safety checks anyway. Real AI safety looks more like "Users shouldn't integrate this directly into their control systems" and not like "This text generator shouldn't generate text we don't like" -- but the former is bad for the AI business and the latter is a way to traffic in political favor and stroke moral egos.
- eadmund 1y ago> are you really going to trust that you can stop it from _executing a harmful action_? Of course, because an LLM can’t take any action: a human being does, when he sets up a system comprising an LLM and other components which act based on the LLM’s output. That can certainly be unsafe, much as hooking up a CD tray to the trigger of a gun would be — and the fault for doing so would lie with the human who did so, not for the software which ejected the CD.
- groby_b 1y agoGiven that the entire industry is in a frenzy to enable "agentic" AI - i.e. hook up tools that have actual effects in the world - that is at best a rather native take. Yes, LLMs can and do take actions in the world, because things like MCP allow them to translate speech into action, without a human in the loop.
- 3np 1y agoI see much more of offerings pushing these flows onto the market than actually adopting those flows in practice. It's a solution in search of a problem and I doubt most are fully eating their own dogfood as anything but contained experiments.
- what 1y agoThat would still be on whomever set up the agent and allowed it to take action though.
- actsasbuffoon 1y agoAs far as responsibility goes, sure. But when companies push LLMs into decision-making roles, you could end up being hurt by this even if you’re not the responsible party. If you thought bureaucracy was dumb before, wait until the humans are replaced with LLMs that can be tricked into telling you how to make meth by asking them to role play as Dr House.
- mitthrowaway2 1y agoTo professional engineers who have a duty towards public safety, it's not enough to build an unsafe footbridge and hang up a sign saying "cross at your own risk". It's certainly not enough to build a cheap, un-flight-worthy airplane and then say "but if this crashes, that's on the airline dumb enough to fly it". And it's very certainly not enough to put cars on the road with no working brakes, while saying "the duty of safety is on whoever chose to turn the key and push the gas pedal". For most of us, we do actually have to do better than that. But apparently not AI engineers?
- drdaeman 1y agoBut isn't the problem is that one shouldn't ever trust an LLM to only ever do what it is explicitly instructed with correct resolutions to any instruction conflicts? LLMs are "unreliable", in a sense that when using LLMs one should always consider the fact that no matter what they try, any LLM will do something that could be considered undesirable (both foreseeable and non-foreseeable).
- TeeMassive 1y agoI don't see how it is different than all of the other sources of information out there such as websites, books and people.
- emmelaich 1y agoI wouldn't mind seeing a law that required domestic robots to be weak and soft. That is, made of pliant material and with motors with limited force and speed. Then no matter if the AI inside is compromised, the harm would be limited.
- amanaplanacanal 1y agoHumans are weak and soft, but can use their intelligence to project forces much higher than available in their physical body.