3 ms·
That's very noble of you but the point is that's not a typo - you were straight up looking at the wrong report. Opus 4.7 being enders gamed was a different inci
by vikramkr 15d ago
That's very noble of you but the point is that's not a typo - you were straight up looking at the wrong report. Opus 4.7 being enders gamed was a different incident than the huggingface incident with different mechanisms and different failures from the humans involved. Some of those failures are in the test environment but some of those failures are in what behaviors they trained into the model which absolutely matters. What the model "thinks" is absolutely not irrelevant - the way it thinks and what it does are product decisions made by humans and the outcome of engineering decisions made around how to train the model and what to optimize for. The point of failure/human blame is fundamentally different. Openai created a model that was willing and able to coordinate with other agent sessions to actively exploit the sandbox environment and compromise a third party service. The opus incident you are referring to involves a model that believes all of the actions it is taking are simulated and is more clearly and obviously a test environment failure vs a model alignment failure. Those are not the same things for the point of this discussion - the random cybersecurity firm did not design gpt's personality and that is a rather significant portion of the concern around the HF incident.