4 ms·
Author here! I think this is reasonable but I have two responses. 1. It's kinda interesting because this is a clear case where the model must be thinking beyon
by josephmiller 4y ago
Author here! I think this is reasonable but I have two responses.
1. It's kinda interesting because this is a clear case where the model must be thinking beyond the next token, whereas in most contexts it's hard to say whether the model thinks ahead at all (although I would guess that it does most of the time).
2. More importantly, the key question here is how it works. We're not surprised that it has this behavior, but we want to understand which exact weights and biases in the network are responsible.
Note also that this is just the introductory sentence and the rest of the article would read exactly the same without it.
> it doesn't seem that surprising that there might be cases where there is a dominant bigram [...] that would trigger an "an" prediction, without actually predicting the following word first
btw I don't really understand what you mean by this. Bigrams can explain the second prediction in a two word pair but not the first.
- defen 4y agoWhat happens if you ask it to complete "I climbed up the pear tree and picked a pear. I climbed up the apple tree and picked a"
- eropple 4y agoI asked this to ChatGPT. Its answer? "Red apple". ...well played.
- bilsbie 4y agoAs a test, during text generation could you change an “a” to an “an” and see if it changes the noun. Or did it already have a noun in mind and it sticks with that.
- tqi 4y agoThanks for responding, and for the cool article! Agree this is pretty tangential to the rest of the article. So (again, I am very much not an expert so please correct me if I'm wrong) I guess my analogy would be if instead of predicting one word at a time, it predicted one letter at a time. At some point, would only be one word that could fit. So if prompted with "SENTE", and it returned "N", that doesn't mean that it's thinking ahead to the "CE" / knows that it is spelling "SENTENCE" already. Is that a correct way to think of it?
- deleted 4y ago[deleted]
- tiarafawn 4y agoI think what the parent was trying to communicate (and what I'm thinking as well) is doubting your premise in 1. ("the model must be thinking beyond the next token"). Rephrase "The model is good at picking the correct article for the word it wants to output next" to "After having picked a specific article, the model is good at picking a follow-up noun that matches the chosen article". Nothing about the second statement seems like an unlikely feat for a model that only predicts one word at a time without any thinking ahead about specific words.
- sebzim4500 4y agoThe prompt is: >I climbed up the pear tree and picked a pear. I climbed up the apple tree and picked The argument made in the article (IMO an extremely convincing one) is that it wouldn't be able to predict the word 'an' except by observing that the word afterwards must be apple. Otherwise why not pick 'a'?
- planb 4y ago> It's kinda interesting because this is a clear case where the model must be thinking beyond the next token I don't see how you get to this conclusion. From all the training data it has seen, "an" is the most probable next word after "I climbed up the tree and picked up". The network does not need to know anything about the apple at this point. Then, the next word is "apple" (with an even higher probability I guess).
- pera 4y agoI don't understand why your conclusion is that "the model must be thinking beyond the next token": the model doesn't need to do that to generate a well-formed sentence because it's not constrained by the size of the sentence.