2 ms·
artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 float
by irthomasthomas 17d ago
artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.
Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.
- Gareth321 17d agoIt's very impressive that DeepSeek 4.1 beats Astra in one benchmark, but I presume you are aware that Astra wins in almost all other benchmarks? These are some of them: https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-gpt-6-astra https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-...