3 ms·
More than other AI moments in the last few years, this feels to me like an event in tech that will be seen retroactively as an important watershed moment. I re
by Willish42 1mo ago
More than other AI moments in the last few years, this feels to me like an event in tech that will be seen retroactively as an important watershed moment.
I really appreciated the negative framing and critical tone of this author's overview. Having also read through the METR report, I fee like frontier labs' pattern of getting PR about how impressive their models are "behind the scenes" has poisoned their ability to take security and reliability seriously. Sure these are new failure modes and the agents operate at a scale that's difficult to combat, but the lack of controls and concern for mitigating these sorts of hacks in the future is crazy to me.
The model for postmortems I was taught which has served me well in my career is thoroughly answering the following:
- what happened / what was the timeline of events?
- how was it mitigated and ultimately resolved?
- what went well?
- what went wrong?
- where did we get "lucky" (meaning it could've gone worse but some arbitrary details about the incident worked out in our favor. usually stuff like "happened during business hours" or "we were already looking at a related thing that brought this to our attention before it was a worse outage")
- (action items) how do we detect, mitigate, and prevent this type of failure in the future?
I really hope OpenAI has done an internal postmortem that answers these questions thoroughly. Most SWEs in the industry have to do such postmortems for much smaller outages with way less impact and risk of societal harm. This is probably another area where regulation and governmental oversight would help curb the risks. What's to stop OpenAI and other frontier labs from an intentional "accidental" attack that results in gaining access to competitors' systems?
I'm also curious what, if anything, Anthropic and Google have done differently to prevent a similar event. I suspect they actually monitored the agents as part of their studies and had better guardrails in their infra for how they set up their harnesses etc. for testing models, particularly when the other guardrails are absent as was the case here.