3 ms·
I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.
by adrianN 2mo ago
I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.
- tsunamifury 2mo agoI'm sorry, we know exactly how LLMs work, this myth that we "dont know how they work" was perpetuated by executives that dont know how they work. We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.
- warkdarrior 2mo agoIf you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example) I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?
- runarberg 2mo agoYou don’t have the computational ability to process as many calculations as a datacenter. You can hardly transpose a 5×5 matrix in your mind, so you won’t be able to do what datacenters do. This is like saying we don‘t know how a car works because a car can beat the best human athletes in 100 meter dash.
- ForHackernews 2mo agoYou can see how an LLM works here https://bbycroft.net/llm https://bbycroft.net/llm they are not magic.
- slopinthebag 2mo agoHow did you generate the next word? Did you first read pretty much every written work ever published, including blog posts, forum posts, books, etc? Learn how to imagine everything as a point in a gigantic abstract space where similar meanings cluster together? How did you manage training with gradient descent? And then did you do a lifetime of matrix multiplication for each token you predicted?
- tsunamifury 2mo agoYes, it turns out that matrix math over a feature space of math works pretty well because unlike poetry or real world work, maths are internally coherent and entirely theoretical.
- thesmtsolver2 2mo agoThis not understanding "understanding". > maths are internally coherent and entirely theoretical Nope. This kind of wish-washy thinking is not what we mean by understanding. https://iep.utm.edu/math-inc/ https://iep.utm.edu/math-inc/
- tsunamifury 2mo agostop with reductionist absurdum.
- gpt5 2mo agoFunny how you said "it turns out" when the whole point is that we don't understand LLMs - we just empirically see what they are good at. Claiming that you understand LLMs is similar to saying that you understand how our biology work because you understand evolution. No - you understand the mechanism behind evolution, but not the complexity it produces.
- tsunamifury 2mo agostop it. This is such reductionist bullshit. By your infinite reductionism definition no one knows how anything works
- airstrike 2mo agoIt turns out being orders of magnitude faster at searching with the aid of a strong verifier is a great way to generate proofs
- pama 2mo agoI agree these are central components, but to avoid oversimplification and the mistaken belief that modern LLMs do a lot of search during inference: If it was so simple, the traditional computer algebra systems would have reached similar breakthroughs when deployed at large supercomputer centers. This didnt happen because the search space is huge. You definitely also need a fancy learning algorithm. Although these ingredients would suffice (depending on what the learning algorithm is), you probably also want to learn in the absense of a strong verifier at every step, to allow building a fuzzy/erratic sense of the search space that can lead to planning/intuition and allow distant jumps in a targetted direction.
- naasking 2mo agoThe search space is far too large for a mere order of magnitude to make any difference at all.
- soiax 2mo agoRight, but we have no clue why, and how the emergent behavior they show works. If we would know that, there would be no need for interpretability research.
- naasking 2mo ago> We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces. This is not what's meant by the statements that we don't know how LLMs work. Explain why LLMs are so good at programming, finding bugs, and developing mathematical proofs. Like, way better than all prior tools specifically designed to be bug finding tools, despite being merely "language models".
- Jensson 2mo agoYou aren't contradicting the person.
- gpt5 2mo agoThey definitely are - the OP claimed that we are reaching "the end of the road for LLMs", based on absolutely no data and some handwaving on pareto distribution. We absolutely don't know enough about LLMs and intelligence to make such a bold (and ridiculous) claim. If anything, all evidence point to the contrary, with new scientific breakthrough achieved across a variety of fields via LLMs. I've been really struggling to understand how the HN community can so boldly claim that LLMs are going to stop improving or not really smart. I just read it as the "denial" stage of the stages of grief that a good portion of this community is in right now (which is understandable).
- runarberg 2mo agoWe know plenty about human cognition, and we know everything about how LLMs work. True we don’t know anything about intelligence but that is because “intelligence” is it self a fraught and vague term, and we haven’t (and perhaps never will) settled on what it means exactly.
- adrianN 2mo agoI’m certainly no expert in the field but to my knowledge a lot of the LLM science is empirical. I’m not aware of a theory that lets us predict what architecture and what number of parameters is needed to solve a particular set of problems.
- pama 2mo agoI wish we knew everything about how LLMs work! We only know very basic elements related to their construction and traning dynamics, and pretty much every major question we would like to address still has unknown or vague heuristic answers. This is expected for such a young field of study. In physics, we know the Schrodinger equation, but we dont know everything about how the world works or how to create new materials even though we know that these materials are composed by atoms and we can simulate small collections of them. In cell biology, we know the sequences that make up the DNA of a cell and we approach the time we can build minimal synthetic cells with pieces we understand, but we only scratch the surface of our level of understanding of how the cells actually work and new discoveries are added every day. In biology at large we still keep finding new types of tubes inside human brains—not sure what you mean by plenty, but we certainly have an extremely limited understanding of human cognition compared to what we might have in 50 years from now. It is not just anout LLMs and intelligence—I would like us to be able to answer practical questions about how LLMs work in order to improve general or specialized LLMs even faster than today. We “know” about scaling in an empirical sense, and it certainly has a long way to go, but it does not feel close to a complete understanding.
- runarberg 2mo agoBoth you and your sibling are approaching LLMs like it is some sort of science. If you do that there is no wonder you have a lot of unanswered questions. LLMs are not a science, they are applied statistics. Making predictions to evaluate hypothesis and constructing theories around the hyperparameters of LLMs is no different then making predictions to evaluate hypothesis and constructing theories around the configurations of Nuclear Power Plants, the latter of course being applied physics.