3 ms·
GPT-5.5-Cyber has already at least hit if not surpassed Mythos capability in cyber tasks. The only reason they're holding back is because once its out everyone
by 827a 4mo ago
GPT-5.5-Cyber has already at least hit if not surpassed Mythos capability in cyber tasks. The only reason they're holding back is because once its out everyone would realize that its capabilities were a step change in March, but are not anymore, yet it costs significantly more and is much slower.
- john_strinlai 4mo agohow did you go about assessing this?
- jansan 4mo agoSo you believe one marketing department more than the other?
- NitpickLawyer 4mo agoThe brits have a step-based benchmark that they use for this - https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5-5-cyber-capabilities https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5... They seem pretty close, in both average and "best run" scores. And, in a highly verifiable domain, "best run" or pass@n is what you're looking for.
- aesthesia 4mo agoWorth looking at the followup post that evaluates the current version of Mythos, which solves one of the main tasks that GPT-5.5-Cyber does not. https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber...
- 827a 4mo agoI believe the correct way to interpret AISI’s findings is that both Mythos and 5.5-Cyber are capable of solving their full benchmark (the only two models that can); Mythos does it with fewer tokens and more consistently. Two things of note: 5.5-Cyber is likely to be substantially cheaper than Mythos, given it is priced around Opus. Additionally: AISI has never tested OpenAI’s best public model and actual Mythos competitor: 5.5-Pro.
- chis 4mo agoBut GPT-5.5-Cyber is also not released publicly?