3 ms·
No one spend ~10 mil USD using gpt 5.1 to solve a millenium prize problem with similar solutions in the training data. So we have no reference for whether 5.1,
by mjburgess 19d ago
No one spend ~10 mil USD using gpt 5.1 to solve a millenium prize problem with similar solutions in the training data. So we have no reference for whether 5.1, with that training data and that amount of moeny, could likewise produce such a solution.
For each generation of advancement the "AI psychosis" of the previous wave wears off. Those who believed 4.6's reasoning was an accurate account of its behaviour; those who believed it had goals and solved useful problems reliably; and so on -- now, attribute only these things to Fable. And no doubt when Fable 6 comes along, it will be only v. 6 that does that.
We have seen fairly marginal progress in LLM reliability and performance since the meaningful start of the high-inference/high-reasoning harness era. It just takes people a few model version bumps to break out of the addiction loop to realise this.
At some point progress will stall entirely, and a couple years after that the spell will break and people will stop treating LLMs "as AI" in the wide-eyed sense, and start treating them as unreliable tools that map Text->Text -- as they do now with earlier model versions.
- usef- 19d agoThere's definitely hype, but the newer models can undoubtedly do things the older ones couldn't. I had multiple long-term issues that earlier models couldn't solve (after repeated attempts) that fable did in 1 shot. On small models too, the differences in what I can trust them with has dramatically changed compared to a few months ago. Unless you breathlessly never touched the limitatioms of earlier models, it's very clear that the wall of limitations has been moving outward. Note also when a new benchmark is released, older models do worse on it than recent models, despite none of them being trained to the benchmark.
- mjburgess 19d agoSure, because those businesses collect training data from users who are working on those problems. Perhaps this will never saturate and frontier users will always provide data to fill last generations gaps. My sense is the economics of that are going to collapse. It's currently extremely expensive to be on this endless retrain and inference cycle in order just to bake in additional marginal features. Maybe, maybe not. However I don't personally see anything other than 'one more leap', which might in any case arise from better integration with harnesses. I can foresee a step change due to harness reinforcement -- but other than that long mild refinements that are very expensive to acquire
- usef- 19d agoWhat tasks are you trying it with? It sounds like whatever you're doing isn't saturating the existing intelligence