4 ms·
it is what I do to solve hard problems through. easy stuff happens by itself, but with a system large enough you need a scratchpad and a rubber duck.
by sznio 1mo ago
it is what I do to solve hard problems through.
easy stuff happens by itself, but with a system large enough you need a scratchpad and a rubber duck.
- dgellow 1mo agoOne thing about the reasoning is that models are trained to generate a chain of thoughts, but it doesn’t have to be correct, accurate, or reflect the underlying logic of the LLM. It’s the same problem we have with the output, it is something plausible, but not that reliable
- anon291 1mo agoI do the same thing in my head. There is no underlying logic to an llm. Logic is an external construct alien to human like forms of reasoning.
- trhway 1mo ago>but it doesn’t have to be correct, accurate, or ... why we don't do GAN here, ie. second model verifying correctness/accuracy/etc. ?
- naasking 1mo agoYes, both the output should be "milestones" of sorts, like lemmas and theorems in math. Important plateaus that serve as a launching pad to the next phase. Regurgitating every thought potentially degrades signal:noise ratio.
- anon291 1mo agoThe hidden states of the tokens likely contain more semantic information than can be extracted by the final projection into token space.
- bee_rider 1mo agoActually, how does chain of thought work? Is the LLM actually creating the tokens and then re-reading them, or is the there still a full hidden state under the hood and then the UI just prints that projection?
- sznio 1mo agoCan't tell you what's happening in a closed model, but in case of open ones it's just a text stream, same "take all previous tokens, compute next token" mechanism applies. Thinking vs Response is just a state change like between a system message and a user message. Closed models probably do the same thing internally. What is shown externally is different though: you get a summary of the chain of thought, not the thoughts itself. This is done to prevent distillation. The latest look we had at a frontier chain of thought is probably in the Huggingface incident report - I haven't actually read it yet but I saw the BlackHat talk, and it included some snippets. The thoughts look like they are approaching neuralese. The words are still understandable but the grammar is weird, simplified. In comparison, Qwen 3.8 27b thinks in valid English.
- someguynamedq 1mo agoRight but it's meant to be smarter than you