3 ms·
Easy to claim things as wrong without prodiving any facts. Can I try? > But when people bring him up, they're of course not generally citing his anger Wrong.
by iLoveOncall 25d ago
Easy to claim things as wrong without prodiving any facts.
Can I try?
> But when people bring him up, they're of course not generally citing his anger
Wrong.
> Google has been increasing the relative priority of revenue over the user experience over time
Wrong.
> I'm curious what people do after being on the wrong side
Wrong.
Some of the claims categorized as "wrong" are also completely true, such as training hitting diminishing returns. New models are barely an improvement and most people I know stuck on Opus 4.6 over any newer one for example.
Exact same thing for the claim "the fact we're running out of high quality training data and we're hitting the walls of scaling laws, in the training paradigm, these models aren't getting better. What we're seeing today is pretty much what they're always gonna be like".
If anything, model performance has regressed in actual use (i.e. not benchmarks) for the past half a year.
- aesthesia 25d ago> Some of the claims categorized as "wrong" are also completely true, such as training hitting diminishing returns. New models are barely an improvement and most people I know stuck on Opus 4.6 over any newer one for example. OK, but the first instance of a claim of diminishing returns was in February 2024, when GPT-4 was the best model available. Do you really think improvement since then has been minimal?
- iLoveOncall 25d agoI personally think improvement has been minimal since ChatGPT was released actually.
- minimaxir 25d agoThat's not defensible.
- iLoveOncall 24d agoIt absolutely is. What has improved isn't the models, it's the harnesses. Give GPT-3.5 a 1M context window and a modern harness, and you won't see any meaningful difference with Opus 5. It's a bit hard to try with such old models, but for example I use Opus 5 / Fable at work and Sonnet 4.5 at home (because it's free via Amazon Q), and there's absolutely 0 difference in performance. None. Obviously 4.5 is only a year old, not 3, but try with any older model that has a decent context window and you'll get the same results. In fact I'll go further than this and say that models are currently regressing. Opus 5 is much much worse than Opus 4.6 for example, and it's clear that Anthropic (at least - I don't use OpenAI models much) is just tokenmaxing rather than optimizing for performance.
- frozenseven 24d ago>Give GPT-3.5 a 1M context window and a modern harness, and you won't see any meaningful difference with Opus 5. >models are currently regressing. Opus 5 is much much worse than Opus 4.6 I'm in sheer awe at these takes. Literally beyond parody.
- iLoveOncall 24d agoStriking argument.
- aesthesia 24d agoBenchmarks are far from everything, but I would love to see the outcome of an experiment benchmarking GPT-4o (which is one of the earlier models with a >100k context window) against GPT-5.6 or Opus 5 in modern harnesses.
- marcosdumay 25d agoYes, since 2024 the improvement has happened only in a few contexts, and the most famous¹ LLMs have also regressed in many contexts. 1 - Their numbers have also exploded, so I have no idea of any general rule.