3 ms·
Coming from a more traditional stats/ML background, I try to view "thinking" traces as a way to explore the search space without getting caught in a local maxim
by perrygeo 1mo ago
Coming from a more traditional stats/ML background, I try to view "thinking" traces as a way to explore the search space without getting caught in a local maximum.
A better analogy for me is annealing; you can't cool metal down instantly or the result is brittle. You must cool down gradually, which allows the molecules to arrange into more durable structures. Random but controlled.
In the same way, thinking traces are testing out all sorts of novel connections between tokens ("But wait..", "Actually,.."). Like a highly divergent branching mind map that gets pruned over time, rather than settling directly into the initial answer.
Now consider human cognition. We're constantly diverging, daydreaming, playing "what if" scenarios and measuring up those ideas against our internal objective functions (proxy for reality) to see which ideas stick. Not too dissimilar. But it's hard to call what we do "thinking" either - it's the default mode network wandering.
- rightbyte 1mo agoIsn't search a strange analogy for how weight terms propagate?
- infinitebit 1mo agothis seems like more anthropomorphizing. even if it has similar results often enough to be useful, next token prediction is not searching a solution space. you even end with a paragraph saying “it’s not too dissimilar from what we do”. how is that not anthropomorphizing? and if what we do isn’t “thinking” then what is or ever has been?
- perrygeo 1mo ago> next token prediction is not searching a solution space. Interesting take. Next token prediction (via the attention mechanism) is a "walk" through the token embedding space. Searching the solution space is what it does, mathematically. It's how we take tokens x context length possible combinations and prune them to converge on viable answers so quickly. Does it look like search at inference time? No. With given weights, a given prompt, and a given random seed, you get the exact same answer. There's not much searching happening at inference... The key is that most of that space is searched at training time. The weights implicitly prune the search space, blocking off or make certain token combinations effectively impossible. It's easy to think "we're just applying weights at inference time" without considering all the pre-work that's done to prune that search space. Which is exactly why "thinking" traces (and randomization) are useful! They bust out of any local optima created by too-tightly-constrained models or system prompts. It's both useful and technically correct to speak of the process as a high-dimension combinatorial search.