4 ms·
> Maybe I am using them wrong but I actually find reasoning/chain-of-thought models worse at some things You are not using them wrong. Sonnet-3.5, famously not
by maeil 2y ago
> Maybe I am using them wrong but I actually find reasoning/chain-of-thought models worse at some things
You are not using them wrong. Sonnet-3.5, famously not a reasoning model, is still easily the best generalist model that isn't enormously expensive and slow (o1). R1 is at times better but also much less consistent, and almost unusably slow.
This is based on a wide-range of tasks, both coding, non-coding objective, as well as clearly subjective non-STEM, often run in parallel across 4+ models for comparison purposes. Both standard chat completion as well as agentic.
o3-mini has been a big disappointment, sometimes being good at one-shot coding but being much worse than Sonnet at everything else. It feels very rushed as a response to Deepseek.
Any kind of benchmark that had ever been quoted by an LLM provider in their own reports is completely worthless.
I think the idea is that since o1 is after all demonstrably SOTA on the large majority of real-world tasks, that is the way to go, because they feel like they can make it cheaper/smaller/faster while keeping that performance.
Whether that actually holds remains to be seen.