2 ms·
We burned 11.7B tokens to find the best cyber AI model
- LoganDark 2mo agoOf course, the model trained most recently does the best job. Do we know what DeepSeek V4 Pro 0813's knowledge cutoff is? For all we know, it's simply working in parallel with how the vulnerabilities were first discovered, even if the solutions themselves weren't trained in. Surprised to see Grok 4.6 performing so well. Shame I can't use it privately -- I would never risk getting my Twitter account banned for that
- exceptione 2mo agoI don't see Fable, but I guess Fable would refuse to work anyway on cyber problems(?) But why would Opus 5 work then?
- LoganDark 2mo agoThe Cyber Verification Program only applies to Opus and Sonnet: https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet https://support.claude.com/en/articles/14604842-real-time-cy...
- k4roshi 2mo agoHow was this so urgent that it required sleepless nights for the team?!
- LoganDark 2mo agoIt just means they really love their job :)
- joker99 2mo agoSince the “magic“ is in the harness: can anyone recommend a good OSS cyber harness? I’ve been experimenting a bunch at $work and for large heterogeneous code bases, just using codex or claude seems to work better than experimenting with tools like code graph/graphql to save on tokens
- colenikol2 1mo agoAnd you could just do it for free: https://nonconfirmed.com/app/ai-tools/ https://nonconfirmed.com/app/ai-tools/