3 ms·
How much is the process of a human child in grade school working through a 3-digit x 3-digit multiplication problem (123 * 456) like a GRPO model with thinking
by dataviz1000 2mo ago
How much is the process of a human child in grade school working through a 3-digit x 3-digit multiplication problem (123 * 456) like a GRPO model with thinking tokens doing the same?
Humans are not born being able to achieve that. It is learned behavior. You and everyone else will remember their teacher saying, "Check your work!" Both the human child and the model work through multiplication problems using the same technique, using the distributive property. They both try to get a reward. For the human child, it is a sense of someone commending them for correctly solving the problem, a reward that probably yields some type of positive dopamine or serotonin feedback loop.
The model solving the problem will have a lower error rate if the first series of tokens created is followed by a series of validation tokens that are subsequently followed by error-correction tokens if there is an error!!!
Maybe it is thinking. Maybe it is remembering to validate and check the work and then remembering to fix the error. For the model trained with RL, why did tokens associated with validation towards the middle of a stream of tokens yield much better results? DeepSeek proved with R1-Zero that a model will learn to verify and correct itself from RL alone with no supervised fine tuning (SFT) teacher ever showing it how. The only reason DeepSeek used SFT was to clean up the reasoning tokens to be human readable. [0] When constrained by SFT, the models will use the double meaning of words -- polysemy -- to satisfy being human-readable while also carrying meaning for what they are working on.
Different people think differently. I watched a viral video of some ~11-year-old child talking to his mom or dad about a stream of a voice in his head. He discovered for the first time that he has a stream of thought. When he goes to school and solves a long multiplication problem, like the stream of tokens from the model, that voice will say to itself (him), "Check your work!"
That is a case of the stream of thought as words being aware of the stream of thoughts as words. Self awareness is a different conversation.
What I think is happening is that the child's stream of thought while solving a multiplication problem in school is likely very similar to an AI model's stream of tokens solving a multiplication problem. And they both were learned. The mechanics are very different, yet, the analogy is apt.
[0] https://huggingface.co/chutesai/DeepSeek-R1-NextN/blob/main/README.md https://huggingface.co/chutesai/DeepSeek-R1-NextN/blob/main/...
- randomImmigrant 2mo agoHow much is the process of a human child in grade school working through a 3-digit x 3-digit multiplication problem (123 * 456) like a GRPO model with thinking tokens doing the same? Very little, if you bother to give the biology of the child at least a cursory glance. Let’s take a short peek: 1. Assuming this is normal grade school, and inflicts math upon children earlier in the day, this is somewhere between 7 and 10/11 am, let’s say? At this point, depending on the age, gender, and maturity of the child, every neuron in their brain involved in math is likely off their midday peak in cognitive function. If we move the class to later, a different subset of students will be at the peak. As far as I am aware, GRPO models do not have such internal temporal rhythms driving their behavior that will shape their performance. 2. How well a given child performs will depend on how hungry they are. But not deterministically. If you trivially think each child is like a computer, you may think the rich kid who had a breakfast buffet before coming to school will do better than the half-starved child of a janitor, but that child might mind the lesson with greater intensity. Or not. It’s not something you can pre-calculate with any certainty. While chip to chip variability is certainly known, I’m yet to hear of a chip deciding to do math better and faster than its fellow chips to prove a point. Or to do significantly worse because it’s distracted by the bird on the window sill. What you are noticing is that there are limited ways to solve a 3x3 digit multiplication. Humans, having standardized the process, have now found a way to record it and plug it into correctly translated signal so the same accurate result can be had without using our own minds in the moment. But where I’d not remotely be shocked if a kid from an uncontacted tribe figured out 3 digit multiplication to keep track of his stone collection, I’d be highly shocked if an H100 that was dumped in the trash by accident somehow figured out anything at all. In fact, if it manage to move any of its electrons around on its own, it would be a certified miracle. And then we could talk about there being real similarity even though the specific atomic composition is different.*