4 ms·
OpenAI and Anthropic both have currently safety teams that look for misbehavior in their models (and to some extent, voluntarily disclose what they find to the
by ashdksnndck 1mo ago
OpenAI and Anthropic both have currently safety teams that look for misbehavior in their models (and to some extent, voluntarily disclose what they find to the public). Going forward, it would be hard for them to argue they don’t know their models do stuff like this.
- junon 1mo agoYes, which is the grey area. "Can / might do" vs "they trained it to do that explicitly" is, I believe, the grey area - whether or not they're the same thing. Intent matters for a lot of this - and "intent" is a pretty strong, well discussed legal term.
- andai 1mo agoYour honor, my LLM spun the turbines real fast, as a practical joke!
- junon 1mo agoIt's not your intent to use a tool that can cause real world damage, to actually do such damage. It was your negligence in that case. A different legal concept than intent.
- andai 1mo agoTo clarify I was referring to Stuxnet here (and the current wave of critical infrastructure hacks, which now have "haha whoops the matmul went a bit funny!" as plausible deniability).
- willy_k 1mo ago> knowingly Intent or negligence.
- weird-eye-issue 1mo agoBut you could also use this to argue in the other way to say that they are using due care and therefore not negligent
- ashdksnndck 1mo agoWell, if you’re writing reports saying “we know our model only decides to commit felonies 0.001% of time which we judge to good enough to deploy” … I’m not sure that gets you off the hook the hook for the felonies.
- s1artibartfast 1mo agoIt works for gun or even car manufacturers. They know that some sales will be used for crime.