3 ms·
According to people who have access to Mythos, it is slightly worse than GPT-5.5-xhigh. At least for security tasks. Hold on, I think this claim needs some har
by throwa356262 5mo ago
According to people who have access to Mythos, it is slightly worse than GPT-5.5-xhigh. At least for security tasks.
Hold on, I think this claim needs some hard data. Here you go gentlemen:
https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5-5-cyber-capabilities https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...
- ACCount37 5mo agoThat claim keeps contradicted hard by other parties, who say Mythos beats 5.5 resoundingly on both autonomous search and discovery and creation of complex exploit chains. There might be a harness difference, but also, this CTF-type benchmark might not capture the capability difference fully.
- nimchimpsky 5mo ago[dead]
- aesthesia 5mo agoSee the later post testing a newer Mythos checkpoint, though: https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber...
- throwa356262 5mo agoFair enough