4 ms·
Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.
by clhodapp 2mo ago
Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.
- cyanydeez 2mo agoI assume theyre searching the local gradient to see if theres a better descent before proceeding.
- c0_0p_ 2mo agoI don't think there's anything like that going on. They just word vomit into a secondary area, and then there is an internal prompt that says "clean this up and summarize for the user".
- Groxx 2mo agoLess "internal prompt" and more "they are trained to summarize after a </think> token"
- astrange 2mo agoThe training methods try not to apply any particular rules to the contents of the thinking text. That's called "optimization pressure on CoT" and is thought to reduce safety by inducing the model to lie (or stop clearly printing its intentions) in the thinking text.
- eigenspace 2mo agoLLMs dont do gradient descent to generate tokens. They are trained by gradient descent, but inference doesnt involve it.
- forgotTheLast 2mo agoThat's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens.
- clhodapp 2mo agoYup. You basically just need something for probability to push off of
- throw310822 2mo agoIt's also interesting because in humans the existence of "Aha!" moments that are not preceded by or are only loosely related to a chain of thought is taken as the proof of the fundamental mystery and irreproducibility of human intelligence. Now the same argument is made to deny that LLMs actually think. Go figure.
- clhodapp 2mo agoIt feels apparent to me that LLMs don't do what is colloquially thought of as thinking. What is less apparent is that humans do.
- throw310822 2mo ago> What is less apparent is that humans do. This seems indeed one obvious hole in the argument of the paper. There is no indication whatsoever that human thinking process is more reliable than LLMs intermediate tokens. Which doesn't make our thinking useless, as messy as it might be. We reorder and explain it after the fact.