8 ms·
It's an interesting review but I really dislike this type of techno-utopian determinism: "When models inevitably improve..." Says who? How is it inevitable? Wha
by SupremumLimit 1y ago
It's an interesting review but I really dislike this type of techno-utopian determinism: "When models inevitably improve..." Says who? How is it inevitable? What if they've actually reached their limits by now?
- Dylan16807 1y agoModels are improving every day. People are figuring out thousands of different optimizations to training and to hardware efficiency. The idea that right now in early June 2025 is when improvement stops beggars belief. We might be approaching a limit, but that's going to be a sigmoid curve, not a sudden halt in advancement.
- deadbabe 1y ago5 years ago a person would be blown away by today’s LLMs. But people today will merely say “cool” at whatever LLMs are in use 5 years from now. Or maybe not even that.
- dingnuts 1y ago5 years ago GPT2 was already outputting largely coherent speech, there's been progress but it's not all that shocking
- tptacek 1y agoMost of the developers I know personally who have been radicalized by coding agents, it happened within the past 9 months. It does not feel like we are in a phase of predictable boring improvement.
- keybored 1y agoRadicalized? Going with the flow and wishes of the people who are driving AI is the opposite of that. To have their minds changed drastically, sure..
- tptacek 1y agoSorry I have no idea what you're trying to say here.
- lcnPylGDnU4H9OF 1y ago> very different from the usual or traditional https://www.merriam-webster.com/dictionary/radical https://www.merriam-webster.com/dictionary/radical Deciding that AI is going nowhere to suddenly deciding that coding agents are how they will work going forward is a radical change. That is what they meant.
- keybored 1y agoDid you miss my second paragraph? https://www.merriam-webster.com/dictionary/paragraph https://www.merriam-webster.com/dictionary/paragraph
- Dylan16807 1y agoCan you explain exactly what you meant by your second paragraph? The ambiguity is why you got that reply. If your second paragraph makes that reply irrelevant, are you saying the meaning was "Your use of 'radicalized' is technically correct but I still think you shouldn't have used it here"?
- dwaltrip 1y agoBold prediction…
- sitkack 1y agoIt is copium that it will suddenly stop and the world they knew before will return. ChatGPT came out in Nov 2022. Attention Was All There Was in 2017, we were already 5 years in the past. Or 5 years of research to catch up to, and then from 2022 to now ... papers and research have been increasing exponentially. Even in if SOTA models were frozen, we still have years of research to apply and optimize in various ways.
- BoorishBears 1y agoI think it's equally copium that people keep assuming we're just going to compound our way into intelligence that generalizes enough to stop us from handholding the AI, as much as I'd genuinely enjoy that future. Lately I spend all day post-training models for my product, and I want to say 99% of the research specific to LLMs doesn't reproduce and/or matter once you actually dig in. We're getting exponentially more papers on the topics and they're getting worse on average. Every day there's a new paper claiming an X% gain by post-training some ancient 8B parameter model and comparing it to a bunch of other ancient models after they've overfitted on the public dataset of a given benchmark and given the model a best of 5. And benchmarks won't ever show it, but even ChatGPT 3.5-Turbo has better general world knowledge than a lot models people consider "frontier" models today because post-training makes it easy to cover up those gaps with very impressive one-prompt outputs and strong benchmark scores. - It feels like things are getting stuck in a local maxima: we are making forward progress, the models are useful and getting more useful, but the future people are envisioning takes reaching a completely different goal post that I'm not at all convinced we're making exponential progress towards. There maybe exponential number of techniques claiming to be ground breaking, but what has actually unlocked new capabilities that can't just as easily be attributed to how much more focused post-training has become on coding and math? Test time compute feels like the only one and we're already seeing the cracks form in terms of its effect on hallucinations, and there's a clear ceiling for the performance the current iteration unlocks as all these models are converging on pretty similar performance after just a few model releases.
- deleted 1y ago[deleted]
- rxtexit 1y ago
- a2128 1y agoI think at this point we're reaching more incremental updates, which can score higher on some benchmarks but then simultaneously behave worse with real-world prompts, most especially if they were prompt engineered for a specific model. I recall Google updating their Flash model on their API with no way to revert to the old one and it caused a lot of people to complain that everything they've built is no longer working because the model is just behaving differently than when they wrote all the prompts.
- whbrown 1y agoIsn't it quite possible they replaced that Flash model with a distilled version, saving money rather than increasing quality? This just speaks to the value of open-weights more than anything.
- Sevii 1y agoModels have improved significantly over the last 3 months. Yet people have been saying 'What if they've actually reached their limits by now?' for pushing 3 years.
- greyadept 1y agoFor me, improvement means no hallucination, but that only seems to have gotten worse and I'm interested to find out whether it's actually solvable at all.
- dymk 1y agoAll the benchmarks would disagree with you
- thuuuomas 1y agoToday’s public benchmarks are yesterday’s training data.
- BoorishBears 1y agoThe benchmarks also claim random 32B parameter models beat Claude 4 at coding, so we know just how much they matter. It should be obvious to anyone who with a cursory interest in model training, you can't trust benchmarks unless they're fully private black-boxes. If you can get even a hint of the shape of the questions on a benchmark, it's trivial to synthesize massive amounts of data that help you beat the benchmark. And given the nature of funding right now, you're almost silly not to do it: it's not cheating, it's "demonstrably improving your performance at the downstream task"
- tptacek 1y agoWhy do you care about hallucination for coding problems? You're in an agent loop; the compiler is ground truth. If the LLM hallucinates, the agent just iterates. You don't even see it unless you make the mistake of looking closely.
- 1y ago
- groby_b 1y agoIt is "inevitable" in the sense that in 99% of the cases, tomorrow is just like yesterday. LLMs have been continually improving for years now. The surprising thing would be them not improving further. And if you follow the research even remotely, you know they'll improve for a while, because not all of the breakthroughs have landed in commercial models yet. It's not "techno-utopian determinism". It's a clearly visible trajectory. Meanwhile, if they didn't improve, it wouldn't make a significant change to the overall observations. It's picking a minor nit. The observation that strict prompt adherence plus prompt archival could shift how we program is both true, and it's a phenomenon we observed several times in the past. Nobody keeps the assembly output from the compiler around anymore, either. There's definitely valid criticism to the passage, and it's overly optimistic - in that most non-trivial prompts are still underspecified and have multiple possible implementations, not all correct. That's both a more useful criticism, and not tied to LLM improvements at all.
- double0jimb0 1y agoAre there places that follow the research that speak to the layperson?
- sumedh 1y agoMore compute mean more faster processing, more context.
- QuantumNoodle 1y agoWhat is ironic, if we buy in to the theory that AI will write majority of the code in the next 5-10 years, what is it going to train on after? ITSELF? Seems this theoretic trajectory of "will inevitably get better" is is only true if humans are producing quality training data. The quality of code LLMs create is very well proportionate on how mature and ubiquitous the langues/projects are.
- solarwindy 1y agoI think you neatly summarise why the current pre-trained LLM paradigm is a dead end. If these models were really capable of artificial reasoning and learning, they wouldn’t need more training data at all. If they could learn like a human junior does, and actually progress to being a senior, then I really could believe that we’ll all be out of a job—but they just do not.