3 ms·
For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least. That being said, I currently prefer Sol / Astra to O
by jasonjmcghee 13d ago
For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.
That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
- vessenes 13d agoI was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
- vintermann 13d agoIt's not just about benchmaxxing. Sincerely targeting those long-autonomy benchmarks is questionable in the first place, because naturally it drives the model to assume more and more about what you want.
- svachalek 13d agoThe target market for frontier models is CEOs who want to lay off entire departments of their company. So the long autonomy benchmarks would seem to be sending exactly the right signal.
- vintermann 12d agoThey're still going to have to communicate with the bots replacing those departments they lay off, or they're going to have a bad time.