3 ms·
Nah, if Anthropic can do it, so can OpenAI: > After finding this incident, we broadened our search to roughly 481 million transcripts—an intentionally wide net
by comp_throw7 22d ago
Nah, if Anthropic can do it, so can OpenAI:
> After finding this incident, we broadened our search to roughly 481 million transcripts—an intentionally wide net, consisting of all transcripts from our Frontier Red Team, many non-cyber evaluations, reinforcement learning (RL) environments, subagent logs, and more. We performed a first-stage scan of this group of transcripts for signs of internet access, such as public IP addresses and web addresses, and a second-stage scan using Claude to review the 9.2 million transcripts the first stage flagged for escalation. This scan re-identified the four incidents and found no other cases of similar or worse severity.
https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents https://www.anthropic.com/research/alignment-assessment-cybe...