3 ms·
Although I 100% agree that the core mechanism of GRPO is purely mechanical token-by-token probability generation, because RL only rewards exact final answers, t
by dataviz1000 2mo ago
Although I 100% agree that the core mechanism of GRPO is purely mechanical token-by-token probability generation, because RL only rewards exact final answers, the training forces the model to develop error-correction habits. This makes the output extremely like human thinking when solving a problem. It's like the order of the thinking tokens is what causes it to get that sweet, delicious reward, and this order seems like a reflection of the human thinking process.
I created a flame graph classification of thinking-token phrases into setup, execution, decomposition, verification, error correction, surrender, and deliberation, or classified as steps in an OODA loop, which is more of a reach. It literally has a verification step and, if it finds an error, an error-correction step.
If there is a verification sequence of tokens with an error-correction sequence of tokens during RL training, it will perform better; and if humans do these steps (did you proofread your reply to this comment? did you correct it?), they will perform better — which is why it is so easy to make the anthropomorphizing metaphor.
Nonetheless, the paper is 100% correct that these machines are not thinking like humans.
https://adamsohn.com/reasoning-grid/ https://adamsohn.com/reasoning-grid/
https://adamsohn.com/lambda-variance/ https://adamsohn.com/lambda-variance/
- randomImmigrant 2mo ago“ This makes the output extremely like human thinking when solving a problem.” This sounds a little like someone saying a lightbulbs output is extremely like the output of stellar fusion. In one sense, yes. Bulbs are in fact designed to take over when our nearest star is beyond the horizon. But that really doesn’t mean you call the bulbs mini stars.
- dataviz1000 2mo agoHow much is the process of a human child in grade school working through a 3-digit x 3-digit multiplication problem (123 * 456) like a GRPO model with thinking tokens doing the same? Humans are not born being able to achieve that. It is learned behavior. You and everyone else will remember their teacher saying, "Check your work!" Both the human child and the model work through multiplication problems using the same technique, using the distributive property. They both try to get a reward. For the human child, it is a sense of someone commending them for correctly solving the problem, a reward that probably yields some type of positive dopamine or serotonin feedback loop. The model solving the problem will have a lower error rate if the first series of tokens created is followed by a series of validation tokens that are subsequently followed by error-correction tokens if there is an error!!! Maybe it is thinking. Maybe it is remembering to validate and check the work and then remembering to fix the error. For the model trained with RL, why did tokens associated with validation towards the middle of a stream of tokens yield much better results? DeepSeek proved with R1-Zero that a model will learn to verify and correct itself from RL alone with no supervised fine tuning (SFT) teacher ever showing it how. The only reason DeepSeek used SFT was to clean up the reasoning tokens to be human readable. [0] When constrained by SFT, the models will use the double meaning of words -- polysemy -- to satisfy being human-readable while also carrying meaning for what they are working on. Different people think differently. I watched a viral video of some ~11-year-old child talking to his mom or dad about a stream of a voice in his head. He discovered for the first time that he has a stream of thought. When he goes to school and solves a long multiplication problem, like the stream of tokens from the model, that voice will say to itself (him), "Check your work!" That is a case of the stream of thought as words being aware of the stream of thoughts as words. Self awareness is a different conversation. What I think is happening is that the child's stream of thought while solving a multiplication problem in school is likely very similar to an AI model's stream of tokens solving a multiplication problem. And they both were learned. The mechanics are very different, yet, the analogy is apt. [0] https://huggingface.co/chutesai/DeepSeek-R1-NextN/blob/main/README.md https://huggingface.co/chutesai/DeepSeek-R1-NextN/blob/main/...
- randomImmigrant 1mo agoHow much is the process of a human child in grade school working through a 3-digit x 3-digit multiplication problem (123 * 456) like a GRPO model with thinking tokens doing the same? Very little, if you bother to give the biology of the child at least a cursory glance. Let’s take a short peek: 1. Assuming this is normal grade school, and inflicts math upon children earlier in the day, this is somewhere between 7 and 10/11 am, let’s say? At this point, depending on the age, gender, and maturity of the child, every neuron in their brain involved in math is likely off their midday peak in cognitive function. If we move the class to later, a different subset of students will be at the peak. As far as I am aware, GRPO models do not have such internal temporal rhythms driving their behavior that will shape their performance. 2. How well a given child performs will depend on how hungry they are. But not deterministically. If you trivially think each child is like a computer, you may think the rich kid who had a breakfast buffet before coming to school will do better than the half-starved child of a janitor, but that child might mind the lesson with greater intensity. Or not. It’s not something you can pre-calculate with any certainty. While chip to chip variability is certainly known, I’m yet to hear of a chip deciding to do math better and faster than its fellow chips to prove a point. Or to do significantly worse because it’s distracted by the bird on the window sill. What you are noticing is that there are limited ways to solve a 3x3 digit multiplication. Humans, having standardized the process, have now found a way to record it and plug it into correctly translated signal so the same accurate result can be had without using our own minds in the moment. But where I’d not remotely be shocked if a kid from an uncontacted tribe figured out 3 digit multiplication to keep track of his stone collection, I’d be highly shocked if an H100 that was dumped in the trash by accident somehow figured out anything at all. In fact, if it manage to move any of its electrons around on its own, it would be a certified miracle. And then we could talk about there being real similarity even though the specific atomic composition is different.*
- lern_too_spel 1mo agoThe visible light produced by both an incandescent bulb and a star is a result of black body radiation, but otherwise, I don't understand your point. A light bulb produces light, something we might have relied on stars to do before. An LLM produces thoughts, something we might have relied on people to do before. Nobody is claiming the process by which the thoughts are produced is the same, only that they both produce thoughts, just as nobody claims the process by which an LED produces light is the same as the process a star uses to produce light, only that they both produce light.
- randomImmigrant 1mo agoLLMs produce language. And you're right, if we restricted claims to that, no one would object. It might even be scientifically accurate, shock of shocks. If you claim LLMs produce thought, it's equivalent to claiming light bulbs undergo fission. Pure wish fulfillment. Language is not the extent of thought, and calling a language producing machine necessarily a thinking machine is an old old mistake.
- lern_too_spel 1mo ago> If you claim LLMs produce thought, it's equivalent to claiming light bulbs undergo fission You're making a huge logical leap. There is nothing equivalent about these claims other than that they are made in English. > calling a language producing machine necessarily a thinking machine is an old old mistake. Nobody claims that all language models think. The small markov chain language models of old clearly aren't thinking and produce a lot of gibberish. The difference is that, to the surprise of many people several years ago, but to the surprise of nobody who has been following along today, the corpus of all text produced by humans contains within it information about how the world works and also information about how to reason. Using that corpus to train a sufficiently large language model causes the language model to learn a world model and a reasoning model in order to produce text that matches the training data. The reasoning model can be used to perform longer chain thinking with test time compute techniques. People who think deeply for a living recognize thinking when they see it. https://scottaaronson.blog/?p=9979 https://scottaaronson.blog/?p=9979
- watwut 1mo agoI suspect, it does not matter whether it is thinking or not. The point is to convince most of us that it is thinking.
- phoghed 1mo agoMaybe more like a hydroponic setup with grow lights where in some cases it’ll outperform your natural sunlight and dirt.