20 ms·
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
- roschdal 3y agoNext Language models teaching themselves to think, then kill humans, based on crawled russian website with secret AI instructions.
- raidicy 3y agoAlthough this is obviously satirical hyperbole dataset poisoning is real and will be underappreciated until the 1st catastrophic example of it happening occurs.
- DFHippie 3y agoMany years ago now I wrote my kids a very simple chatbot to play with. You'd type in a phrase. It would tokenize it, adding start and stop tokens, then update it's token transition probabilities, using the two preceding tokens to pick the next one. It would then generate a response from these probabilities. The data poisoning began immediately. Because "poop" was such a funny word, they quickly taught it that the most probable token after any bigram was "poop". No humans were killed, but two small kids were amused for an hour or so.
- raidicy 3y agoMy condolences for your models' poisoning. It sounds like a real crappy way to go :?
- astrange 3y agoIt isn't really all that real, of course untrusted text contains things that aren't true. So don't trust it.
- dcrimp 3y agoI had this thought the other day that the whole chain of thought reasoning pattern contributing to improved performance in LLM-based systems seems to sit parallel to Kahneman's two-system model of the mind that he covers in 'Thinking, Fast and Slow'. Haven't read it in a few years, but I recall the book suggests that we use one 'System 1' in our brains primarily for low-effort, low computation thinking - like 1+1=? or "the sky is ____". It then suggests that we use a 'System 2' for deliberate, conscious, high-cognitive tasks. Dense multiplication, reasoning problems, working with tools - generally just decision-making. Anything that requires focus or brain power. Our brain escalates tasks from S1 to S2 if they feel complex or dangerous. Maybe I'm being too cute, but it feels like critique that "LLMs aren't intelligent because they are stochastic parrots" is an observation that they are only equipped to use their 'System 1'. When we prompt an LLM to think step-by-step, we allow it a workspace to write down it's thoughts which it can then consider in it's next token prediction, a rudimentary System 2, like a deliberation sandbox. We do a similar thing when we engage our System 2 - we hold a diorama of the world in the front of our mind, where we simulate what the environment will do if we proceed with a given action - what our friend might respond to what we say, how the sheet steel might bend to a force, how the code might break, how the tyres might grip. And we use that simulation to explore a tree of possibilities and decide an action that rewards us the most. I'm no expert, but this paper seems to recognise a similar framework to the above. Perhaps a recurrent deliberation/simulation mechanism will make it's way into models in the future, especially the action models we are seeing in robotics.
- OJFord 3y agoI'm currently reading it for the first time, completely coincidentally/not for this reason, and on a few occasions I've thought 'Gosh that's just like' or 'analogous to' or 'brilliant description of that problem' for LLMs/generative AI or some aspect of it. I wish I could recall some examples.
- machiaweliczny 3y agoIt’s a bit over my head for now but seems like GFlowNets are tackling this problem a bit.
- adlpz 3y agoAny relation to OpenAI's rumored Q* (i.e. q-star) model? Authors of this paper don't seem affiliated. Just a name coincidence?
- HarHarVeryFunny 3y agoI was thinking the same. The STaR paper this is an extension of came out in 2022, so at least possible this is what q-star is based on too, but maybe with Q standing for something else.
- smusamashah 3y agoI think it's just a play on the same hyped up term.
- anon291 3y agoSo it seems 'obvious' to me that a network about 50 layers deep (for example) can only reason about symbolic questions for 50 'steps' (in quotes because it's not a step as we think about it). It only seems there's more complexity because it's 50 steps in one or more learned subspaces that the model has been trained in (which might mean the model can accomplish more than one 'human step' in its 'step'). Humans (well intelligent humans at least) seem able to obviously reason beyond those steps, but we all know it requires real thinking and deliberation and perhaps a notepad to be able to do that. It's quite something to, for example, expect ChatGPT to be able to correctly do 4 digit multiplications without any thought or recourse to 'paper' when very few human beings can do that.
- blackbear_ 3y agoThis paper does indeed follow your intuition to investigate the limits of transformers on compositional tasks (i.e., those that require multi-step reasoning, including your multiplication example): https://arxiv.org/abs/2305.18654 https://arxiv.org/abs/2305.18654 > Our empirical findings suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations' performance can rapidly decay with increased task complexity.
- anon291 3y agoAh good... This is definitely a research path I've been looking into. Great to see someone else has already gone there!
- visarga 3y agoMaybe the Skill Mix paper is relevant here. They define a list of 100 skills, and then randomly sample tuples of n skills (usually less than 6) and generate a test example using those skills. Apparently only GPT-4 (at the time of the paper) was able to compose 5 skills, the other models just 3 or 2. Beyond 5 skills even GPT-4 was doing much worse. The interesting finding of the paper is that GPT-4 couldn't have seen all the (topic, skill-tuple) combinations in the training set. If you have 10,000 examples on a topic, and use 5 out of 100 skills, you would need 100^5 training examples to cover all combinations. In conclusion GPT-4 generalizes to new skill combinations, thus it is not a stochastic parrot. https://arxiv.org/abs/2310.17567 https://arxiv.org/abs/2310.17567
- FeepingCreature 3y agoHere we go!! I've been waiting years for them to try this. Let's see how it does when scaled up to GPT-3/4 level. This might be the missing piece to AGI.
- parthianshotgun 3y agoThe missing piece is unknowable
- Cthulhu_ 3y agoWe'll likely reconstruct what the missing piece was in hindsight, but it's very probable there's no one missing piece. Just like human evolution.
- digging 3y agoUntil it's been found, you mean?
- sroussey 3y agoMaybe even then too!
- arendtio 3y agoI am not convinced there even is a missing piece. I mean, LLMs are being used very differently compared to how traditional AI programs were written. Combining both worlds might be all that is needed. I would not be surprised that when we have general artificial intelligences, we will see, that advancing LLMs wasn't necessary.
- 082349872349872 3y agoEdsger Dijkstra had a precise english style; even though his mother tongue was Dutch, I find he made better use of English than many native speakers. In one of the EWD's, he reminisced that, as children, they were taught to never begin to speak a sentence unless they already knew how they were going to finish it. I'd bet these two observations have a causal connection.
- MattPalmer1086 3y agoThat actually sounds like hell to me, a complete absence of spontaneity and being in the moment. I used to obsessively try to figure out what to say before I said it. I am socially awkward, and it did not help at all. I love writing because it is asynchronous and I can figure things out precisely and edit my thoughts. But in social situations it is a complete hindrance.
- fennecbutt 3y agoUnfortunately from experience that just gives enough of a delay that you get talked over in a group setting and never get a chance to speak anyway.
- ricardobeat 3y agoIs that even possible, or just hyperbole? I'd bet the latter. I wouldn't be surprised if some people are able to fully unravel entire paragraphs of conversation in their head in a couple of seconds, but that's not something you could teach to children in general.
- mannykannot 3y agoI don't think it is feasible, at least for conversation, but as an aspirational goal for children, along the lines of "put your toys away when you've finished playing with them", it is not a bad one. It's not unusual for me to think I know how I am going to end a sentence, but then find that I can't get there.
- h34t 3y agoin Dutch (and German) the verb often goes at the end of a sentence, so the advice is rather practical.
- QuantumG 3y agoWe're done for!
- pawnty 3y agoThis is the missing piece to train AI which has the ability to reason. There are so many tasks whose answers are known but reason steps are missing. With this method, we can use less annotated data the reach the ability. The interesting part(I imagine): the generated thought could be hard for human to understand while it is still way more helpful to get the correct answer! If that happens, we have created something more intelligent than ourselves.
- silent_cal 3y agoNeural networks do not think
- adlpz 3y agoDo neurons think? Do a bunch of neurons? Is this semantics?
- empath-nirvana 3y agobasically this: https://en.wikipedia.org/wiki/Sorites_paradox https://en.wikipedia.org/wiki/Sorites_paradox One neuron doesn't think. Three neurons don't think. Billions of neurons think. Somewhere between one neuron and billions of neurons, thinking starts happening. Probably also true for neural networks. The main problem is that people throw around terms like: "Thought", "Intelligence", "Will", "Reasoning", "Knowledge", "Consciousness", etc like they are very well defined and well understood terms and they very much are not.
- silent_cal 3y agoBillions of neurons don't think, people do.
- empath-nirvana 3y ago...with what?
- silent_cal 3y agoWith their minds
- adlpz 3y agoMy point precisely. Those are all vague terms. Saying that "neural nerworks do not think" is as meaningless as any equivalent (or opposite) statement on any other system including any number of neurons, a whole brain or a person. It's all semantics.
- YetAnotherNick 3y agoAnother RL paper with terrible baseline. They used 0 shot non instruction tuned Mistral for GSM8k which has very specific way of output. They got 11% accuracy after improving it, while few shot prompting achieves 37%[1]. GPT 4 could get ~97% with prompting. [1]: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
- hiddencost 3y agoFwiw if they're serious scientists, taking a known method and baseline and improving it is good science. Extensions to get state of the art are probably possible, but their goal is to measure just the impact of their change in a simple setting. Let the engineers do the munged system combinations and get SoTA.
- YetAnotherNick 3y agoI am not talking about SoTA. I am talking about deliberate poor baseline. GSM8k consists of two things: solving the problem and getting the output format correct. Getting the output format corrects gives 30% accuracy for the same model where they got 11%. SoTA is 97%.
- lionkor 3y agoThis is purely anecdotal, and I try to keep it to myself but its very difficult when at least half of the HN homepage is AI related: LLMs like ChatGPT do so utterly terribly at any non-trivial job I throw at it that I seriously consider people who use it daily to either be straight up incompetent, or maybe their domain is so trivial that the LLM actually does well. From asking LLMs to solve a highly difficult async C++ parallelism problem, to german language specifics, it just fucks up at a fundamental level. I understand that LLMs cannot solve these issues and why, but then I do not understand the heavy focus on AI by so many tech people. Is day to day programming job so trivial that LLMs do a good job, while at the same time being too difficult for you to do it yourself? I really, really want to know exactly what the use case is. Do people just throw simple problems at it to validate their own preconceived notion of how cool and useful LLMs are? Whats the deal?
- rplnt 3y agoEvery other query I've given to ChatGPT came up with an utterly wrong answer. Followup always yielded "sorry, I made an obvious mistake, here's another wrong answer". Confident and stupid is a very bad combination.
- jollyllama 3y agoThere are plenty of jobs where people have to complete various tasks that are outside of their domain or otherwise tedious on a daily basis. For example, plenty of devs have to set up or change the configuration of remote hosts. Some LLMs are pretty good at generating devops scripts to speed up this work.
- orzig 3y agoExactly. Example: maybe 1% of the code I generate is bash. I used to try to memorize patterns, but of the top 20 I'd use each less than once per year. Now, instead of that 1% taking 5% of my time, it takes 2%. It's all "simple stuff", and I can verify it instantly. I have ~10 similar use cases. So it hasn't revolutionized my life, but it's been well worth $20/mo ChatGPT Plus and $3/mo API calls.
- 3y ago
- archibaldJ 3y agoThis looks really interesting; any possibility the researchers will release some code soon ?
- iAkashPaul 3y agoBase Mistral 7B is hardly suitable for the evaluations, even one team at Intel tried to pull a fast one with NeuralChat in the exact same way https://huggingface.co/Intel/neural-chat-7b-v3#quantitative-analyses https://huggingface.co/Intel/neural-chat-7b-v3#quantitative-...
- kjqgqkejbfefn 3y agoThis is basically what I tried this morning at the prompt level (awful results), but the sketchy idea I had in mind went further by introducing control-flow "meta-tokens" to help the LLM renavigate its context. In this perspective the context would be rethought as a self-editing structured mind-map, with the linear aspect of the context at a time T standing for the execution trace of the exploration of this mind-map so far. Some of those meta-tokens would be able to have side effects on the context, to highlight, give structure, summarize, forget and so on, some of its parts. This could allow for native structured output without using a syntactic format such as json, programmatic constructs in the style of LMQL, implementing memory, etc. The goal: not just to give logical/reasoning abilities to a LLM, but to give it the means to come up with its own cognitive architecture. Implementing structured output (using a <label name="stuff">...</label> token) to also implement memory/scratchpads, would also bring inspectability of those cognitive structures for free. Of course I have no idea how to implement this (I'm a ML tourist).
- lawlessone 3y agoIf it is doing this , is it still a language model? or also a thought model?
- aaroninsf 3y agoObservation: "expertise" (hence "reflex") is the learning of the nonlinear solution space that can be inferred from initial conditions. Conjecture: models which engage in self-training on the solutions they derive will get to something that looks a bit like bootstrapping when you squint. Lemma: there's a nice opportunity for cloud-hosted model SaaS to offer discounts for actionable feedback on the quality of their output, so as to drive this retraining. Idle comment: I'd use the language of REM sleep and the idea of "memory consolidation" for this. Most of the premises of model training can be extended to the level of reasoned solutions, rather than tokens.
- thesz 3y agoThey do not cite [1], a paper on (learned) variable computation in RNN, applied to language modeling, that predates their work by almost 8 years. [1] https://openreview.net/pdf?id=S1LVSrcge https://openreview.net/pdf?id=S1LVSrcge Microsoft also had something alike at that time, but for image recognition: a CNN at input and then varable computation at classification.
- itissid 3y ago> Much of the meaning of text is hidden between the lines: without understanding why statements appear in a document, a reader has only a shallow understanding. This seems not true when I and most people I know read things. I would argue that we almost always have a world model and know some reasons why these statements are appearing in a book. If I was reading Fluid Dynamics textbook, I may not understand the math, but I know why those statements appear; They are mathematical statements to help you learn the theory or whatever and they follow a pattern to teach you important concepts. Like for instance concepts will build upon older ones Bernoulli's equation is there because law of conservation of energy was there before it, so its there because it assumes I understand the latter..