4 ms·
No, it’s not a threshold. It’s just how the tech works. It’s a next letter guesser. Put in a different set of letters to start, and it’ll guess the next letter
by clysm 1y ago
No, it’s not a threshold. It’s just how the tech works.
It’s a next letter guesser. Put in a different set of letters to start, and it’ll guess the next letters differently.
- Trasmatta 1y agoI think we need to start moving away from this explanation, because the truth is more complex. Anthropic's own research showed that Claude does actually "plan ahead", beyond the next token. https://www.anthropic.com/research/tracing-thoughts-language-model https://www.anthropic.com/research/tracing-thoughts-language... > Instead, we found that Claude plans ahead. Before starting the second line, it began "thinking" of potential on-topic words that would rhyme with "grab it". Then, with these plans in mind, it writes a line to end with the planned word.
- cmiles74 1y agoIt reads to me like they compare the output of different prompts and somehow reach the conclusion that Claude is generating more than one token and "planning" ahead. They leave out how this works. My guess is that they have Claude generate a set of candidate outputs and the Claude chooses the "best" candidate and returns that. I agree this improves the usefulness of the output but I don't think this is a fundamentally different thing from "guessing the next token". UPDATE: I read the paper and I was being overly generous. It's still just guessing the next token as it always has. This "multi-hop reasoning" is really just another way of talking about the relationships between tokens.
- Trasmatta 1y agoThat's not the methodology they used. They're actually inspecting Claude's internal state and suppression certain concepts, or replacing them with others. The paper goes into more detail. The "planning" happens further in advance than "the next token".
- cmiles74 1y agoOkay, I read the paper. I see what they are saying but I strongly disagree that the model is "thinking". They have highlighted that relationships between words is complicated, which we already knew. They also point out that some words are related to other words which are related to other words which, again, we already knew. Lastly they used their model (not Claude) to change the weights associated with some words, thus changing the output to meet their predictions, which I agree is very interesting. Interpreting the relationship between words as "multi-hop reasoning" is more about changing the words we use to talk about things and less about fundamental changes in the way LLMs work. It's still doing the same thing it did two years ago (although much faster and better). It's guessing the next token.
- Trasmatta 1y agoI said "planning ahead", not "thinking". It's clearly doing more than only predicting the very next token.
- therealpygon 1y agoThey have written multiple papers on the subject, so there isn’t much need for you to guess incorrectly what they did.
- ceh123 1y agoI'm not sure if this really says the truth is more complex? It is still doing next-token prediction, but it's prediction method is sufficiently complicated in terms of conditional probabilities that it recognizes that if you need to rhyme, you need to get to some future state, which then impacts the probabilities of the intermediate states. At least in my view it's still inherently a next-token predictor, just with really good conditional probability understandings.
- jermaustin1 1y agoBut then so are we? We are just predicting the next word we are saying, are we not? Even when you add thoughts behind it (sure some people think differently - be it without an inner monologue, or be it just in colors and sounds and shapes, etc), but that "reasoning" is still going into the act of coming up with the next word we are speaking/writing.
- thomastjeffery 1y agoWe are really only what we understand ourselves to be? We must have a pretty great understanding of that thing we can't explain then.
- hadlock 1y agoHumans and LLMs are built differently, it seems disingenuous to think we both use the same methods to arrive at the same general conclusion. I can inherently understand some proofs of pythagorean's theorem but an LLM might apply different ones for various reasons. But the output/result is still the same. If a next token generator run in parallel can generate a performant relational database that doesn't directly imply I am also a next token generator.
- wetpaws 1y ago[dead]
- spookie 1y agoThis type of response always irks me. It shows that we, computer scientists, think of ourselves as experts on anything. Even though biological machines are well outside our expertise. We should stop repeating things we don't understand.
- dontlikeyoueith 1y ago> Anthropic's own research showed that Claude does actually "plan ahead", beyond the next token. For a very vacuous sense of "plan ahead", sure. By that logic, a basic Markov-chain with beam search plans ahead too.