9 ms·
Procedural knowledge in pretraining drives reasoning in large language models
- largbae 2y agoIs this conclusion similar to my layman's understanding of AlphaGo vs AlphaZero? That human procedural knowledge helps ML training to a point, and from there on becomes a limitation?
- dinfinity 2y agoNo. They're saying that the model they analyzed used mainly information on _how_ to solve math problems from its training data, rather than documents that contained the answers to the (identical) math problems: > "We investigate which data influence the model’s produced reasoning traces and how those data relate to the specific problems being addressed. Are models simply ‘retrieving’ answers from previously seen pretraining data and reassembling them, or are they employing a more robust strategy for generalisation?" > "When we characterise the top ranked documents for the reasoning questions qualitatively, we confirm that the influential documents often contain procedural knowledge, like demonstrating how to obtain a solution using formulae or code. Our findings indicate that the approach to reasoning the models use is unlike retrieval, and more like a generalisable strategy that synthesises procedural knowledge from documents doing a similar form of reasoning." Example reasoning question: > "Prompt Calculate the answer: (7 - 4) * 7 Think step-by-step."
- spitfire 2y agoWhat I further got from this is the models are learning the methods, but not evaluating themselves along the way. They don’t check for errors. So once they go down a path they can’t properly backtrack. This feels like the ground truth I’ve experienced in LLMs to date.
- spitfire 2y agoI’ll add when I say “learning” I mean memorization. Memorizing on a higher level than facts. I would love to spend the time and see how altering the query alters the reasoning path. How firm is in the path once it’s chosen? A high level approach has the possibility to be very computer efficient.
- Nevermark 2y ago> Memorizing on a higher level than facts Which is not memorization, since memorization is defined by its limits: storing information based on its literal form, as apposed to some higher meaning. It is called generalization. Learning specific examples with a shared memory too small to memorize all the examples, creates a gradient toward a more compact storage form: patterns. Which, unlike memorized examples, are able to generate reasonable guesses for similar but previously unencountered problems. Generalization does not require reasoning, nor is it required for reasoning. But they often complement each other. Where reasoning usually means some kind of flexible application of multiple steps. I.e. a sequence of steps, trying alternative steps, stepping forward to a solution, stepping back from the goal, accumulation of ever larger solved subsets or substeps of the problem, etc.
- NitpickLawyer 2y ago> So once they go down a path they can’t properly backtrack. That's what the specific training in o1 / r1 / qwq are addressing. The model outputs things like "i need to ... > thought 1 > ... > wait that's wrong > i need to go back > thought 2 > ... etc
- sgt101 2y agodrives retrieval of patterns of procedure? I mean - like for arithmetic?
- ijk 2y agoThis would explain the unexpected benefits of training on code.
- strken 2y agoThat sounds interesting, but I'm a layman and don't know anything about it. Can you provide a link? I was able to find https://arxiv.org/abs/2408.10914 https://arxiv.org/abs/2408.10914, but I don't have the context to know whether it's the paper you're talking about.
- MurizS 2y agoI think GP was probably referring to "Scaling Data-Constrained Language Models" (2305.16264) from NeurIPS 2023, which looked first at how to optimally scale LLMs when training data is limited. There is a short section on mixing code (Python) into the training data and the effect this has on performance on e.g. natural language tasks. One of their findings was that training data can be up to 50% code without actually degrading performance, and in some cases (benchmarks like bAbI and WebNLG) with improvements (probably because these tasks have an emphasis on what they call "long-range state tracking capabilities"). For reference: In the Llama 3 technical report (2407.21783), they mention that they ended up using 17% code tokens in their training data.
- eru 2y agoIs the network only trained on the source code, or does it have access to the results of running the code, too?
- YetAnotherNick 2y agoAlso GPT-3.5 was another extreme if I remember correctly. They first trained only on code then they trained on other text. I can't seem to find the source though.
- moffkalast 2y agoThere was an interview with Zuckerberg about how they initially split training llama chat models on purely normal text and codellama on code, but later realized that if they combine the training set they get a model that is better at both tasks than each specialized one was.
- jpcom 2y agoYou mean you need humans to step-by-step solve a problem so a neural net can mimic it? It sounds kinda obvious now that I write it out.
- mattdeboard 2y agoNo. If I'm understanding correctly it means the software is learning how to solve problems in general by ingesting examples of procedural problem-solving.
- jpcom 2y agoYou're close, but there’s an important nuance. The process isn't about "learning how to solve problems in general" in the broad sense. It's more specific: the neural network is trained to mimic the step-by-step process demonstrated by humans solving a specific problem. The distinction is that the software doesn't autonomously derive general problem-solving heuristics from scratch. Instead, it observes examples of how humans solve problems procedurally and uses that to replicate similar reasoning. This is crucial because the step-by-step demonstrations give the model structure and guidance, which is different from learning a generalizable strategy for solving any kind of problem without those examples. In essence, it's like a neural net learning to follow a recipe by watching a chef cook—rather than inventing its own recipes entirely from first principles.
- jebarker 2y ago> In essence, it's like a neural net learning to follow a recipe by watching a chef cook—rather than inventing its own recipes entirely from first principles. Just like how a chef learns
- Retric 2y agoA chef also learns through trial and error not just reading how others have cooked in the past and then copping their motions. This is exemplified by how altitude has a meaningful impact but isn’t discussed for a given recipe.
- semessier 2y agothat resonates - less facts and more reasoning training data. The most low hanging in terms of non synthetic data probably being mathematical proofs. With prolog and the like many alternate reasoning paths could be generated. It's hard to say if these many-path would help in llm training without access to the gigantic machines (it's so unfair) to try it on.
- deleted 2y ago[deleted]
- shermantanktop 2y agoGoing meta a bit: comments so far on this post show diametrically opposing understandings of the paper, which demonstrates just how varied the interpretation of complex text can be. We hold AI to a pretty high standard of correctness, as we should, but humans are not that reliable on matters of fact, let alone on rigor of reasoning.
- sigmoid10 2y agoThis is extremely common in these discussions. Most humans are not that good at reasoning themselves and fall for the same kind of fallacies over and over because of the way they were brought up (their training data so to speak). And yet they somehow think they can argue why or why not LLMs should be able to do the same. If anything, the current limits of these morels show the limits of human cognition which is spread throughout the internet - because this is literally what they learned from. I believe once we achieve a more independent learning (like we've seen glimpses of in the MuZero paper) these models will blow human intelligence out of the water.
- pclmulqdq 2y agoIt's because we can put responsibility on humans to be correct but we can't on computers. Humans given the appropriate incentives are very good at their jobs and there is a path for compensation if they screw up. Computers have neither of these things.
- sigmoid10 2y agoHumans already put a lot of trust in computers not because they can take responsibility but because traditional software can be made very predictable or at least compliant. There are whole industries built around software standards to ensure that. The problem is we don't yet know enough about identifying and patching problems in these models. Once we get something equivalent to MISRA for LLMs to achieve the same level of compliance, there is very little that could still hold them back.
- 2y ago
- ninetyninenine 2y ago>On the one hand, LLMs demonstrate a general ability to solve problems. On the other hand, they show surprising reasoning gaps when compared to humans, casting doubt on the robustness of their generalisation strategies surprised this gets voted up given the surprising amount of users on HN who think LLMs can't reason at all and that the only way to characterize an LLM is through the lens of a next token predictor. Last time I was talking about LLM intelligence someone rudely told me to read up on how LLMs work and that we already know exactly how they work and they're just token predictors.
- ben_w 2y agoThe loudest people seem to be those with the most extreme positions, and that includes on "is ${specific AI} (useless|superhuman) for ${domain}?". Perhaps it's just perception, but perhaps the arguments make them persist, as CGP Grey pointed out: https://www.youtube.com/watch?v=rE3j_RHkqJc https://www.youtube.com/watch?v=rE3j_RHkqJc As I'm in the middle, I get flack from people on both extremes, as I'm outside their (equivalent of or just literally?) Overton window on this subject. Seems like an odd zone to be in for the opinion "this is a useful tool, but I see loads of ways it can go wrong". Makes me wonder what the real common discourse was of looms during the industrial revolution, and not just the modern summary of that era.
- tkgally 2y ago> Makes me wonder what the real common discourse was of looms during the industrial revolution, and not just the modern summary of that era. Interesting question. I did a little searching with help from Claude, ChatGPT, Google, and the Internet Archive. Here are some links to writings from that time: “Thoughts on the use of machines, in the cotton manufacture” (1780) https://archive.org/details/bim_eighteenth-century_thoughts-on-the-use-of-m_1780 https://archive.org/details/bim_eighteenth-century_thoughts-... Excerpt: “How many writers and copiers of books were thrown out of employment, or obliged to change it, by the introduction of printing presses? About ten years ago, when the Spinning Jennies came up, old persons, children, and those who could not easily learn to use the new machines, did suffer, for a while; till families had learned to play into one another's hands, by each taking a different kind of work. But the general benefit, which was received from the machines, very soon silenced all objections. And every sensible man now looks upon them with gratitude and approbation. It is probable, this will be the case in all new inventions.” Kevin Binfield, ed., Writings of the Luddites (1811-1815; 2004) https://ia903409.us.archive.org/16/items/writings-of-the-luddites/Writings%20of%20the%20Luddites.pdf https://ia903409.us.archive.org/16/items/writings-of-the-lud... Robert Owen, Observations on the effect of the manufacturing system (1817) https://archive.org/details/observationsonef00owenrich/page/n5/mode/2up https://archive.org/details/observationsonef00owenrich/page/... William Radcliffe, Origin of the new system of manufacture commonly called power-loom weaving (1828) https://archive.org/details/originofnewsyste0000radc https://archive.org/details/originofnewsyste0000radc “An address to the Glasgow cotton-spinners on the moral bearing of their association” (1838) https://catalogue.nla.gov.au/catalog/6023196 https://catalogue.nla.gov.au/catalog/6023196
- rors 2y agoIt seems obvious to me that LLMs wouldn't be able to find examples of every single problem posed to them in training data. There wouldn't be enough examples for the factual look up needed in an information retrieval style search. I can believe that they're doing some form of extrapolation to create novel solutions to posed problems. It's interesting that this paper doesn't contradict the conclusions of the Apple LLM paper[0], where prompts were corrupted to force the LLM into making errors. I can also believe that LLMs can only make small deviations from existing example solutions in creation of these novel solutions. I hate that we're using the term "reasoning" for this solution generation process. It's a term coined by LLM companies to evoke an almost emotional response on how we talk about this technology. However, it does appear that we are capable of instructing machines to follow a series of steps using natural language, with some degree of ambiguity. That in of itself is a huge stride forward. [0] https://machinelearning.apple.com/research/gsm-symbolic https://machinelearning.apple.com/research/gsm-symbolic
- ucefkh 2y agoTotally, these companies are pushing towards showcasing their AI models as self thinking and reasoning AI while they are just trained of a lot of amount of data in dataset format which they extrapolate to find the right answer. They still can't think outsider their box of datasets
- pfisherman 2y agoI very much agree with the perspective that LLMs are not suited for “reasoning” in the sense of creative problem solving or application of logic. I think that the real potential in this domain is having them act as a sort of “compiler” layer that bridges the gap between natural language - which is imprecise - and formal languages (sql, prolog, python, lean, etc) that are more suited for solving these types of problems. And then maybe synthesizing the results / outputs of the formal language layer. Basically “agents”. That being said, I do think that LLMs are capable of “verbal reasoning” operations. I don’t have a good sense of the boundaries that distinguish the logics - verbal, qualitative, quantitative reasoning. What comes to my mind is the verbal sections of standardized tests.
- 2y ago
- ricardobeat 2y agoDoes this mean LLMs might do better if trained on large amounts of student notes, exams, book reviews and such? That would be incredibly interesting.
- GarnetFloride 2y agoI have wondered that from time to time, why not train an AI system using educational curricula plus some games and play? It might be fascinating to see what comes out using various systems from around the world.
- ilaksh 2y agoThey do train on textbooks.
- jpcom 2y agoYep.
- btilly 2y agoThis is highly relevant to the recent discussion at https://news.ycombinator.com/item?id=42285128 https://news.ycombinator.com/item?id=42285128. Google claims that their use of pretraining is a key requirement for being able to deliver a (slightly) better chip design. And they claim that a responding paper that did not attempt to do pretraining, should have been expected to be well below the state of the art in chip design. Given how important reasoning is for chip design, and given how important pretraining is for driving reasoning in large language models, it is obvious that Google's reasoning is very reasonable. If Google barely beats the state of the art while using pretraining, an attempt that doesn't pretrain should be expected to be well below the current state of the art. And therefore that second attempt's poor performance says nothing about whether Google's results are plausible.
- pfisherman 2y agoI am not an expert in the particular application domain of that article; but I can see why their argument of pre training might be valid. It is not especially controversial to say that pre training neural nets improves few shot learning performance. And I suspect there is an inflection point for every problem where pre trained neural nets yield better few shot learning performance than less data hungry approaches - such as hand crafted features or strong priors. That being said, it seems that the question here is whether that inflection point has been reached in this case.
- andai 2y ago> In the extreme case, a language model answering reasoning questions may rely heavily on retrieval from parametric knowledge influenced by a limited set of documents within its pretraining data. In this scenario, specific documents containing the information to be retrieved (i.e. the reasoning traces) contribute significantly to the model’s output, while many other documents play a minimal role. > Conversely, at the other end of the spectrum, the model may draw from a broad range of documents that are more abstractly related to the question, with each document influencing many different questions similarly, but contributing a relatively small amount to the final output. We propose generalisable reasoning should look like the latter strategy. Isn't it much more impressive if a model can generalize from a single example?
- ScottPowers 2y agothanks
- samirillian 2y agoOkay dumb question, why are the images they generate nightmarish nonsense. Why can’t they procedurally construct a diagram