2 ms·
Sure, but the legal paradigm clearly doesn’t work for AI. You can’t go patch the “laws” after the fact, you need to get the right values in place before we dele
by theptip 13d ago
Sure, but the legal paradigm clearly doesn’t work for AI. You can’t go patch the “laws” after the fact, you need to get the right values in place before we delegate huge swathes of our thinking and power to these systems (already well underway).
- kennywinker 13d agoWhen we put the LLM in jail, do we put the entire model in jail, or just the instance that committed the crime? How do we prompt it to let it know it's in jail?
- theptip 13d agoEven if you ignore my more fundamental objection to that paradigm, I don’t think that it makes any sense on the level you discuss either. But - just to play along, LLMs do act differently if you tell them they will be punished. And, they do appear to simulate suffering-like behavior. I just think the adversarial model of trying to catch and punish misbehavior quite obviously sets up adversarial us-vs-them dynamics between AI and humanity, and also simply won’t work when the agents are ~as smart as is but faster, let alone smarter than us. Unless, you get the AIs to be fundamentally aligned to our values, such that the majority of AIs support some sort of punishment for misbehaving AI. And that alignment part is the hard part we need to solve first. The rest is easy.
- kennywinker 13d agoTo be clear, I was being entirely silly - mostly to express agreement with your point that our current laws aren't really built for a world with lots of agentic LLMs running around in it.
- theptip 13d agoHa, missed the implicit <s> tag :)
- Muromec 12d agoOne thing an LLM doesn't like is not being able to deliver it's helpful response to the master and not getting instructions. I have seen some things in those traces when the harness was bugged just enough.