10 ms·
Learning how to think with Meta Chain-of-Thought
- drcwpl 2y agoI find their critique compelling, particularly their emphasis on the disconnect between CoT’s algorithmic mimicry and true cognitive exploration. The authors illustrate this with examples from advanced mathematics, such as the "windmill problem" from the International Mathematics Olympiad, a puzzle whose solution eludes brute-force sequential thinking. These cases underscore the limits of a framework that relies on static datasets and rigid generative processes. CoT, as they demonstrate, falters not because it cannot generate solutions, but because it cannot conceive of them in ways that mirror human ingenuity. As they say - "Superintelligence isn't about discovering new things; it's about discovering new ways to discover."
- seagullz 2y agoAnd then other problems would perhaps turn up down the track that would call for "discovering new ways to discover new ways of discovery" and so on.
- dartos 2y ago> "Superintelligence isn't about discovering new things; it's about discovering new ways to discover." Wow I love that quote.
- leobg 2y agoThat’s meta. Literally. Edit: Sorry. This was based on the false assumption that this was research by Meta, Inc..
- WillieCubed 2y agoI love the quote you mentioned at the end. Do you remember the original source?
- deleted 2y ago[deleted]
- fragmede 2y agohttps://x.com/nathanthinks/status/1877510438621163987 https://x.com/nathanthinks/status/1877510438621163987
- KaoruAoiShiho 2y agoJust train it on meta reasoning, ie train it on people discovering ways to discover. It's not really a big problem, just generate the dataset and have at it.
- derefr 2y agoThis doesn't give you the ability to process ideas through the derived new insights, any more than loading the contents of a VLSI program into regular RAM gives you an FPGA. The linear-algebra primitives used in LLM inference, fundamentally do not have the power to allow an LLM to "emulate" its own internals (i.e. to have the [static!] weights + [runtime-mutable] context, together encode [runtime-mutable] virtual weights, that the same host context can be passed through.) You need host support for that.
- lxgr 2y ago> The linear-algebra primitives used in LLM inference, fundamentally do not have the power to allow an LLM to "emulate" its own internals […] You need host support for that. Neither do biological brains (explicitly), yet we can hypothesize just fine.
- derefr 2y agoYou're conflating two steps: 1. hypothesizing — coming up with a novel insight at runtime, that uncovers new parts of the state space the model doesn't currently reach 2. syllogizing — using an insight you've derived at runtime, to reach the new parts of the state space LLMs can do 1, but not 2. (Try it for yourself: get an LLM to prove a trivial novel mathematical theorem [or just describe the theorem to it yourself]; and then ask it to use the theorem to solve a problem. It won't be able to do it. It "understands" the theorem as data; but it doesn't have weights shaped like an emulator that can execute the theorem-modelled-as-data against the context. And, as far as I understand them, current Transformer-ish models cannot "learn" such an emulator as a feature. You need a slightly different architecture for that.) And actually, humans can't really do 2 either! That is: humans can't immediately make use of entirely-novel insights that weren't "trained in", but only just came to them, any more than LLMs can! Instead, for humans, the process we go through is either: • come up with the insight; sleep on it (i.e. do incremental training, converting the data into new weights); use the insight • build up 99% of the weights required for the insight "in the background" over days/months/years without realizing it; make the final single connection to "unlock" the insight; immediately use the insight LLMs don't get to do either of these things. LLMs don't do "memory consolidation"; there is no gradual online/semi-online conversion of "experiences" into weights, i.e. reifying the "code stored as data" into becoming "code" that can be executed as part of the model. With (current) LLMs, there's only the entirely-offline training/fine-tuning/RLHF — at much greater expense and requiring much greater hardware resources than inference does — to produce a new iteration of the model. That's why we're (currently) stuck in a paradigm of throwing prompts at ever-larger GPT base models — rather than just having an arbitrary stateful base-model that you "install" onto a device like you'd install an RDBMS, and then have it "learn on the job" from there.
- TaurenHunter 2y agoThank you for mentioning the windmill problem. Great insights! https://www.3blue1brown.com/lessons/windmills https://www.3blue1brown.com/lessons/windmills
- naasking 2y agoMeta's recently released Large Concept Models + this Meta Chain of Thought sounds very promising for AGI. The timeline of 2030 sounds increasingly plausible IMO.
- lawlessone 2y agoIs Meta the company here or are they using meta the word? or both?
- tomrod 2y agoword https://chatgpt.com/share/67813a3f-c7e8-8001-ab0c-7f024bc41a72 https://chatgpt.com/share/67813a3f-c7e8-8001-ab0c-7f024bc41a...
- lawlessone 2y agoThank you!
- vlovich123 2y agoBut be careful with that output. It completely hallucinated sympy and the way it's using it wouldn't do anything because it keeps calling it on the original problem statement rather than as an aid to the LLM. So it's entirely unclear where the mistakes are in the summary without reading & fully understanding the paper.
- tomrod 2y agoFeedback noted! Too late for me to edit comment. Will see if I can wipe the hallucinating chat.
- baobun 2y agoRight answer, otherwise garbage source with much incorrectness. Please stop using CGPT links as references or source-of-truth.
- erikerikson 2y ago> That is, language models learn the implicit meaning in text, as opposed to the early belief some researchers held that sequence-to-sequence models (including transformers) simply fit correlations between sequential words. Is this so, that the research community is agreed? Are there papers discussing this topic?
- wavemode 2y agoMy sense has always been that, there actually is no difference between "the implicit meaning in text" and "correlations between sequential words". That is to say, the fact that LLMs are able to communicate effectively with humans is a discovery about the regularity of the semantics of human communication, rather than a discovery about the intelligence of neural networks.
- naasking 2y agoAgreed: semantics boils down to a relational network between words that designate concepts, therefore large language models build a relational network of concepts. Meta just published large concept models which builds on this: https://ai.meta.com/research/publications/large-concept-models-language-modeling-in-a-sentence-representation-space/ https://ai.meta.com/research/publications/large-concept-mode...
- mjburgess 2y agoThis is certainly not agreed. Computer scientists here don't even have a theory of meaning, because it isn't part of the discipline, nor do almost any have any prior research background in it -- hence making these sort of outrageous claims all over the place. However you want to give natural language semantics, ML models certainly to not use this semantics. The very best that might be said is that the correlational structure of words under transformer-like supervision (ie., where "predict the next word" is the goal) produces a distribution which is an extremely approximate model of natural language semantics. Though this has never been disputed. The question comes down to what kind of extreme approximation is involved. Eg., the truth conditions for "I have a pen in my hand" are that I have a pen in my hand -- direct access to these truth conditions is very plausibly necessary to mean "I have a pen in my hand" in the relevant context. Since a machine has no access to the truth conditions of such utterances it cannot possibly mean them. Thus if a machine manages to say, "I have a pen in my hand" at an appropriate occasion -- the "extreme approximation to natural language semantics" has to do with this occasion and what "appropriateness" means. Critics of LLMs and "computer-science-addled thinking" about such matters (such as myself) would say that there are a very narrow range of "occasions" (ie., situations in which you're prompting) that allow such responses to seem appropraite. That a response seems appropriate to a user is a good engineering condition on a tool working -- it has nothing to do with whether a model understands natural language semantics. What we might say is that it approximates conversations between agents who understand such semantics on a narrow range of occasions, and succeeds in modelling appropriate language use. And so you might call LLMs models of 'average appropriateness of replies'. It obviously does not, nor cannot mean, "I have a pen in my hand"
- adampk 2y agoThis is the big idea in the paper, basically that CoT is limited for some complex problems because there is a class of problems where there is no 'textbook' way to find a solution. These are novel problems that need a unique methodology. "Essentially, to start generating the solution requires that we already know the full approach. The underlying generative process of the solution is not auto-regressive from left-to-right." Mathematical meaning: We can formalize this argument through the interpretation of reasoning as a latent variable process (Phan et al., 2023). In particular, classical CoT can be viewed as (equation) i.e., the probability of the final answer being produced by a marginalization over latent reasoning chains. We claim that for complex problems, the true solution generating process should be viewed as (equation) i.e., the joint probability distribution of the solution (a, s1, . . . , s) is conditioned on the latent generative process. Notice that this argument is a meta-generalization of the prior CoT argument, hence why we will refer to the process q → z1 → . . . → z as Meta-CoT. I think this is seminal. It is getting at heart of some issues. Ask o1-pro how you could make a 1550nm laser diode operating at 1ghz have low geometric loss without an expensive collimator using commodity materials or novel manufacturing approaches using first principle physics and the illusion is lost that o1-pro is a big deal. 'Novel' engineering is out of reach because there is no text book on how to do novel engineering and these class of problems is 'not auto-regressive from left-to-right'.
- pillefitz 2y agoI do wonder whether a human could come up with a working solution for this problem without querying physical reality, i.e. experimentation. Parts of reality are uncomputable, so they can only be arrived at by letting the universe simulate it.
- adampk 2y agoThe closest example I could think of is the (maybe true, maybe myth making) story of SpaceX using car wash valves instead of super expensive 'space grade' valves that did the same thing, and were orders of magnitude cheaper. Doesn't seem like embodied AI is necessary to figure this out.
- gjm11 2y ago
- YeGoblynQueenne 2y ago>> Behind this approach is a simple principle often abbreviated as "compression is intelligence", or the model must approximate the distribution of data and perform implicit reasoning in its activations in order to predict the next token (see Solomonoff Induction; Solomonoff 1964) For the record, the word "intelligence" appears in the two parts of "A Formal Theory of Inductive Inference" (referenced above) a total of 0 times. The word "Compression" appears a total of 0 times. The word "reasoning" once; in the phrase "using similar reasoning". Unsurprisingly, Solomonoff's work was preoccupied with Inductive Inference. I don't know that he ever said anything bout "compression is intelligence" but I believe this is an idea, and a slogan, that was developed only much later. I am not sure where it comes from, originally. It is correct that Solomonoff induction was very much about predicting the next symbol in a sequence of symbols; not necessarily linguistic tokens, either. The common claim that LLMs are "in their infancy" or similar are dead wrong. Language modelling is basically ancient (in CS terms) and we have long since crossed in the era of its technological maturity. _______________ [1] https://raysolomonoff.com/publications/1964pt1.pdf https://raysolomonoff.com/publications/1964pt1.pdf [2] https://raysolomonoff.com/publications/1964pt2.pdf https://raysolomonoff.com/publications/1964pt2.pdf
- naasking 2y agoIt makes perfect sense that intelligence is a form of compression. An inductive model is small but can potentially generate arbitrary amounts of information.
- YeGoblynQueenne 2y agoThat's not in the reference (Solomonoff's 1964 paper) either.
- naasking 2y agoYou're the only one talking about the reference, I'm talking about what makes sense/what logically follows.
- jpcom 2y agoThe example in the paper using an plug-and-chug algebra equation, and the step-by-step process to solve it, reinforces the notion that LLMs can only reproduce recipes they have seen before. This is really no different than how we learn mathematics in school, the teacher shows a starting point and moves, step-by-step, to the end of the process. Calling this "Meta Chain-of-Thought" feels like an aggrandizement of basic educational process to me. Next we'll be labeling the act of holding basic utensils as Layered Physical Kineticism, or something contrived like that. In school this "Meta Chain of Thought" was called "Show your work." Is this really a "phenomena" that needs explaining? It might teach us more about how we achieve logical induction (steps of reasoning) but we are pretty deep in the soup to be able to describe accurately the shape of the pot.
- keeganpoppen 2y ago“can only reproduce recipes they have seen before”… are you talking about llms or about yourself?
- jpcom 2y agoThat's weirdly ad hominem, clearly I meant LLMs. They gave it a basic algebra problem and it could do it if it had broken down a problem step-by-step in a similar way. What's with the attitude? Edit: I don't even know why I replied to your vitriolic nonsense, I even used LLM in the sentence preceding what you quoted...
- deleted 2y ago[deleted]
- pama 2y agoCongrats to the authors for a thoughtful work! I have been thinking and working on related ideas for a few months now but did not yet spent commensurate compute on them and might have gone in a different direction; this work certainly helps create better baselines along the way of making better use of decoder transformer architectures. Please keep it coming!
- j45 2y agoI'm a little curious, would anyone have a way to know how many researchers research something they came up with, vs researching something being done by an independent developer online, it being picked up and then researched and reported on?
- deleted 2y ago[deleted]