2 ms·
> Someone didn't get the memo that for LLMs, tokens are units of thinking. Where do you get this memo ? Seems completely wrong to me. More computation does not
by Rexxar 6mo ago
> Someone didn't get the memo that for LLMs, tokens are units of thinking.
Where do you get this memo ? Seems completely wrong to me. More computation does not translate to more "thinking" if you compute the wrong things (ie things that contribute significantly to the final sentence meaning).
- staminade 6mo agoThat’s why you need filler words that contribute little to the sentence meaning but give it a chance to compute/think. This is part of why humans do the same when speaking.
- jaccola 6mo agoDo you have any evidence at all of this? I know how LLMs are trained and this makes no sense to me. Otherwise you'd just put filler words in every input e.g. instead of: "The square root of 256 is" you'd enter "errr The er square um root errr of 256 errr is" and it would miraculously get better? The model can't differentiate between words you entered and words it generated its self...
- lijok 6mo agoYou’re conflating training and inference
- staminade 6mo agoWhat do you think chain of thought reasoning is doing exactly?
- muzani 6mo agoIt's why it starts with "You're absolutely right!" It's not to flatter the user. It's a cheap way to guide the response in a space where it's utilizing the correction.
- mike_hearn 6mo agoPeople have researched pause tokens for this exact reason.
- dTal 6mo agoThe LLM has no accessible state beyond its own output tokens; each pass generates a single token and does not otherwise communicate with subsequent passes. Therefore all information calculated in a pass must be encoded into the entropy of the output token. If the only output of a thinking pass is a dumb filler word with hardly any entropy, then all the thinking for that filler word is forgotten and cannot be reconstructed.