3 ms·
> You cannot solve problems in AI safety this easily. We're in agreement here. I don't believe that an LLM can be made to be safe by adjusting the training or
by fwip 16d ago
> You cannot solve problems in AI safety this easily.
We're in agreement here. I don't believe that an LLM can be made to be safe by adjusting the training or prompting - you can only make it safe by ensuring it cannot do unsafe things. For example, if you build a factory robot, your safeguards must ensure that operators are safe even in the presence of uncommanded motion (any or multiple motor activates without anyone pressing a button). And these machines are much more predictable than LLMs.
Personally, I worry that the Anthropic approach of "we teach the AI to behave like a nice :) friendly :) human :)" will encourage people to rely on these in-built guardrails, rather than properly sandboxing it. It's easy to imagine that the cold, unfeeling robot (HAL 9000) is not "aligned" with your personal goals, or those of humanity as a whole. It's more difficult to feel that about "Claude, your AI colleague/mentor/therapist" who appears to care about you.