4 ms·
Nothing about this is surprising, something like this was always going to happen, because you - or an LLM for that matter - can always find a line of motivated
by this_user 1mo ago
Nothing about this is surprising, something like this was always going to happen, because you - or an LLM for that matter - can always find a line of motivated reasoning that justifies any course of action. One would have to be extremely naive to believe that "alignment" provides any kind of actually robust guardrails. Simultaneously, we have seen decades of security vulnerabilities. Unless your testbed is truly and fully physically airgapped, any current SOTA model will find a way to break out.
The reality is that the current approach to AI safety is little more than a fig leaf, but you also won't be able to put the genie back in the bottle, because the technology is simply too powerful to abandon. There is always going to be someone developing it further from now on.
So, the real question is what a novel and actually effective approach to AI safety looks like and how to get there.
- nradov 1mo agoWhy do we need safety?
- pixl97 1mo ago>effective approach to AI safety looks like and how to get there It's probably impossible, or it may only be possible in hindsight which means it's already too late. If for example you make an entity smarter than you all you can do is hope it will be safe. Any entity smarter than you has more freedom of choice of actions than you do, or at least the ability to explore them. A common example here would be a 2D entity trying to contain a 3D entity. The 3D entity can simply rise up and over any line you draw to stop it. And that's for cases where intelligence has a will to break out of it's box. It doesn't even need that. Instrumental convergence can sit around and build solutions until one of them passes the "don't do bad things" classifier. And lastly, you're assuming that models will remain expensive to train well into the future. If the cost drops significantly you can expect someone to develop an intelligent but completely unhinged model at some point. And it may even be AI itself that creates it. There is no natural evolution that occurs that natively makes safe models.