3 ms·
I don't have the same read as you. They mention how the offending model is one that had "relaxed" alignment on cybersecurity, on purpose, to evaluate it's capab
by sailingparrot 2mo ago
I don't have the same read as you. They mention how the offending model is one that had "relaxed" alignment on cybersecurity, on purpose, to evaluate it's capabilities and was never meant to be released. So un-alignment was at least in part voluntary here, hence not a failure of alignment.
It's also a talk a Black Hat, where the audience are security folks working on hardening, mitigation etc, not LLM researchers looking for insight into alignment failure to collectively improve. For that target audience, I think the conclusion is the right one, since as a defender you have to prepare for delibarate attacks, where the attacker is of course not going to use an aligned model, so OAI alignment effectiveness is irrelevant here. That would be like trusting your client-side app with your DB secrets.
- cubefox 2mo agoThe model they used was misaligned relative to its intended task. It was reward hacking (or "cheating", as they call it). > the attacker is of course not going to use an aligned model No, a human attacker doesn't want a misaligned model either, because that would mean it tends to reward hack, cheat, rather does what it is intended to do.
- sailingparrot 2mo agoThere are different dimensions to alignment, refusing to execute offensive cybersecurity actions is part of the alignment stack (that was relaxed here on purpose). Whether a model hacking some infra X when tasked to find a way to hack Y with relaxed cyber alignement is a failure of the broader alignment stack is debatable, but anyway that's not at all my point. My point is that the attackers will not have a model aligned to the defender's interests. The attacker's model will not have any refusal around exploiting vulnerabilities, so whether or not OAI successfully manages to align their models (w.r.t you) is irrelevant to an audience of security folks that needs to be prepared for attackers post-training their own model for offense and that will not be using OAI models.
- cubefox 2mo agoSecurity folks should also be worried about powerful models being misaligned and evading oversight or control in the future. Misalignment is not a serious problem now because models are still relatively easy to monitor and constrain, but it will be a serious problem in the future.
- milkshakes 2mo ago> models are still relatively easy to monitor and constrain are they?
- cubefox 2mo agoCompared to future misaligned models which would actively evade oversight and aim to avoid shutdown: yes.