6 ms·
This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is
by otterdude 2mo ago
This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression.
If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence.
I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm
“The seeker after truth is not one who studies the writings of the ancients and, following his natural disposition, puts his trust in them, but rather the one who suspects his faith in them and questions what he gathers from them, the one who submits to argument and demonstration and not the sayings of human beings whose nature is fraught with all kinds of imperfection and deficiency. Thus the duty of the man who investigates the writings of scientists, if learning the truth is his goal, is to make himself an enemy of all that he reads, and, applying his mind to the core and margins of of its content, attack it from every side. he should also suspect himself as he performs his critical examination of it, so that he may avoid falling into either prejudice or leniency.” - ibn al-Haytham
- 0xdeadbeefbabe 2mo agoThe seeker of truth must also hold his breath.
- graemep 2mo agoI am wondering whether the reason he needed to say it was because he was arguing with those who did out their trust in the writings of the ancients.
- scotty79 2mo agoDo you draw that conclusion from the fact that AI surprisingly quickly reaches the end of each ruler we try to measure it with?
- freejazz 2mo agoCan't call it AI like that without discrediting yourself. You mean LLMs?
- otterdude 2mo agoJumping in here, frankly I hate the trend of calling every type of automation intelligence. Most "AI" is really an optimization algorithm in software tools, same as its always been. This really isnt anything new, aside from adding a chatbot / MCP interface to the same tools.
- logicchains 2mo agoTalk about moving the goalposts. Pray tell, exactly what must an LLM do before you're willing to consider it AI? Be specific, otherwise you're just woo-mongering.
- scotty79 2mo agoWhen a Big Killing Robot comes to murder you be sure to always call it BKR and don't discredit yourself by calling it AI.
- freejazz 2mo agoThat's a bit hyperbolic when we're all just posting on HN
- otterdude 2mo agoIts not really that surprising when models are trained on the exams
- astro1234 2mo agoI work in AI evaluation, lots of problems and leakage is an issue as is ecological validity, but they definitely do not explain the progress we see. I think Epoch has the best analysis I’ve seen on evaluation trends; they use IRT to basically model a variety of benchmark difficulties, and then model a capability parameter for each model. This is as robust a sort of “meta-study” of evaluations as I’ve seen and the trend in capabilities show no sign of slowing down. So I think people’s feelings clash with reality, and that’s because releases are more frequent and the jumps between releases are smaller, but the growth in capabilities _over time_ has not changed for the better or worse over a very very long period of time.
- otterdude 2mo agoBenchmarks saturate around 80-90%? This is not "Acing" a test, this is hitting a wall.
- scotty79 2mo agoEven on very small tests a fraction of questions might have wrong answers in the key. If models can't get more than 90% of the benchmark right I think it's a strong indication that they were not trained on the answers and that benchmark itself is messy enough that <10% desired answers might be wrong or misleading.
- astro1234 2mo agoYea this may explain part of it or all of it, it’s likely a case by case kind of thing. Also to respond to the parent comment: benchmarks have a variety of difficulty levels. Humanity’s Last Exam, though now hitting the beginning of a saturation phase with Fable, was long unsaturated while other benchmarks saturated awhile ago. So that’s what I meant by Epoch capability index: using IRT models this effect so that you gather robust signals from variety of benchmark difficulties and can track progress over time as model capabilities have evolved (and so benchmarks have had to evolve to keep up). But yes like I was saying: all benchmarks are problematic, some are useful. Benchmark quality problems abound, so 90% being the true ceiling is not surprising. There may be other factors at play here too, I haven’t studied this problem that deeply to have a good thorough answer to this. But keep in mind there are probably 50,000 benchmarks in the literature and that is not a joke number. A crapload of noise in that signal but it’s not all noise.
- adrianN 2mo agoI really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.
- tsunamifury 2mo agoI'm sorry, we know exactly how LLMs work, this myth that we "dont know how they work" was perpetuated by executives that dont know how they work. We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.
- warkdarrior 2mo agoIf you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example) I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?
- runarberg 2mo agoYou don’t have the computational ability to process as many calculations as a datacenter. You can hardly transpose a 5×5 matrix in your mind, so you won’t be able to do what datacenters do. This is like saying we don‘t know how a car works because a car can beat the best human athletes in 100 meter dash.
- ForHackernews 2mo agoYou can see how an LLM works here https://bbycroft.net/llm https://bbycroft.net/llm they are not magic.
- slopinthebag 2mo agoHow did you generate the next word? Did you first read pretty much every written work ever published, including blog posts, forum posts, books, etc? Learn how to imagine everything as a point in a gigantic abstract space where similar meanings cluster together? How did you manage training with gradient descent? And then did you do a lifetime of matrix multiplication for each token you predicted?
- logicchains 2mo ago>This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. It's just inadequate benchmarks. Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.
- Jensson 2mo agoYes, and good senior software engineer is ahead of fable, but benchmarks can't capture that either. We already know from testing humans that test scores don't correlate that well with how effective a person is at work. Same applies here, we just aren't that great at making good tests.
- otterdude 2mo agoIf most models were getting 100% on the test it would be an inadequate benchmarks. What were seeing is all models failing to ace these tests. "Benchmark Saturation" is term that promotes lowering the bar.
- deleted 2mo ago[deleted]
- Planktonne 2mo ago> Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this. I have seen people claim the exact opposite. If there was so much progress, then you wouldn't have endless disagreements with people championing their own favourite model as the strongest.
- ofjcihen 2mo agoThe two times I tried to use fable I had it attempt something I had already had Opus 4.6 do with no issues. It blasted through 10% of my weekly allowance on a 100$ a month sub and produced something broken and nonsensical.
- heaney-555 2mo ago>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.
- Jensson 2mo agoThere is high rate of progress in specific domains, not high rate of progress in generalness. The models haven't gotten generally smarter, for things they didn't focus on the models are just as bad as a year ago.
- naasking 2mo agoI don't think that's correct. Increasing parameter count increases capabilities in all domains per the scaling laws. Models are larger than they were years ago, so capabilities in all domains must necessarily be better. This doesn't even account for better training data, which has also much improved.
- runarberg 2mo agoAnd sales of disco records were up 400% for the year ending 1976. If these trends continue...
- malfist 2mo agoCan you quantify this rate of progress? Because someone always comes around and says there's been exponential progress in the past <short timeline> every time someone complains that models just aren't very good. Both can't be true
- Certhas 2mo agoIn research we are still seeing massive jumps. Subjects that LLMs were completely useless for half a year ago are now definitely in scope. And there are benchmarks that cleanly separate the SOTA models: https://epoch.ai/MirrorCode https://epoch.ai/MirrorCode Saturation of benchmarks is a property of benchmarks just as much as of the models.
- 2mo ago
- hiddencost 2mo agoWeird moment for this take. We're seeing some of the fastest and most impressive progress ever right now. Frontier labs have categorically different & better set ups for evaluation, they're fine. It's work but it's not a crisis.
- malfist 2mo agoI've heard that every week of every month for the past three years. And yet, ask an LLM about a seahorse emoji and see what happens.
- CollinEMac 2mo agoThe seahorse emoji thing appears to be fixed actually. (Sonnet 5)
- ACCount37 2mo agoAnd it held true every week of every month for the past three years. AI progress is screaming forward at a breakneck pace.
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- antisthenes 2mo agoIt's also important not to put too much faith into ancient sayings and aphorisms. As a civilization, we are currently brushing up against the physics of efficiency. In many areas we have achieved close to what is theoretically possible, based on physics. Such was not the case for the majority of human existence. The body of research a.k.a. "writings of the ancients" is now insurmountably higher than it would have been during the time of ibn al-Haytham, when any kind of writing at all was scarce and literacy was low.
- slopinthebag 2mo agoIdk about end of the road, I’m sure they can squeeze out some more performance by curating even more data and doing even more RL. But I would bet that pretty much all of the improvement we’ve seen over the last year with coding has come from RL, not from the models becoming particularly stronger. And this makes sense, if models grow sublinearly with compute. And it seems like they do.
- gr_norm 2mo agoIt seems pretty obvious from the steep 'intelligence' drop-off on out-of-distribution tasks that the performance improvement is from throwing untold tens of billions at RL. There are legions of highly skilled people employed solely to feed the RL loop. Evidently effective, but there's an unmistakable feeling this won't ultimately be the way forward.
- freeone3000 2mo agoWell, why not? Won’t it get “good enough” at every task eventually?
- slopinthebag 2mo agoWhy would that be the default assumption?
- freeone3000 2mo agoBecause they’ve gotten good enough at lots of other things, and the RL keeps improving them, so enough RL should make them good enough at the focus areas.
- slopinthebag 2mo agoI've gotten pretty strong in the gym, my bench has improved to two plates. I see no reason why it won't continue to improve until I can bench my house.
- grim_io 2mo agoAssuming you are completely correct about the 80/20 rule, we have evidently not yet reached that 80%. Who can say when it will be achieved? The ceiling is glass, we have to touch it to know where it is.
- huflungdung 2mo ago[dead]
- holoduke 2mo agoEven without getting better trained models and only speed increase, the output would be dramatically better. An LLM or non llms that is a billion times faster than now would be so insanely strong in many areas.
- DiscourseFan 2mo agoI think LLMs will continue to improve in the capabilities which they are demonstrably good at, but there are many things which they are not good at which it is not cost effective or meaningful to improve, and in these areas we will not consider them “intelligent,” in the same way that we don’t consider computers “intelligent” but do find them very good at doing wrote calculations.
- naasking 2mo ago> but there are many things which they are not good at which it is not cost effective or meaningful to improve Can you name a few such things so I can keep an eye on them in the coming years?
- DiscourseFan 2mo agoNot explicitly, no, but its pretty obvious if you work in the industry and aren't blinded by AGI hype
- naasking 2mo agoThat's just vibes then. How is that supposed to be convincing?
- DiscourseFan 2mo agoUnless you are working for a competing firm or have a specific clause in your contract, you can literally just sign up for AI projects as a contractor and start contributing. You will see and understand everything after a few months. Nobody is hiding this knowledge really.
- naasking 2mo agoI use LLMs regularly for real work. I understand the current limitations, but you're talking about existing limitations that will never be overcome, because they cannot be overcome. I'm just asking for a couple of specific examples, and your evasions are increasingly suspicious.
- pama 2mo agoThe paper suggests the opposite of your first statement. The benchmarks become useless because the successive models keep saturating them.