4 ms·
So does the program while True: print('a') GPT-n (not chat) generates continuations to text in which the probability distribution for the next out
by JoshuaDavid 3y ago
So does the program
while True:
print('a')
GPT-n (not chat) generates continuations to text in which the probability distribution for the next output it produces matches the distribution of outputs after similar text in the training corpus, by most metrics you might examine. Trivial examples of those metrics include word and n-gram frequencies, but also it turns out that at sufficiently low loss that looks like "given a couple of input texts and a context which affords insightful commentary, produce insightful commentary at the base rate".
There are, of course, caveats. Notably:
- ChatGPT has its own stuff going on around chat tuning, tool use, tuning to be less likely to produce outputs OpenAI doesn't want, etc
- The statistical regularity thing is not magic. If it requires more than n_layers steps of computation to determine the sensible next output, the model will not be able to do that. I think the canonical example here is usually having the model complete something like "'d072c916029965a7676da4244160c413e31bc8a0' == sha1('I saw it on hn'); '265149165fcb742a900a44b8f123885dc6ac5d12' == sha1('" and then having the model brute-force sha1 -- obviously it's not going to be able to do that.
- The model generates text sequentially, one token at a time. This is importantly not the process by which most text in the training corpus was written. So in the cases where earlier text importantly depends on text which was written earlier temporally, but which occurs later in the string, the model will be likely to make mistakes (where a "mistake" is "writing text which is statistically surprising, relative to the training corpus").
- DiggyJohnson 3y agoIs there not a way for me to express the crucial difference between non-language character repetition (or repetition of any string) and the ability or the ability to interpret and respond to human language in the same language? I just feel like we're not in a position to even begin understanding our disagreement until you at least recognize the question or point I'm trying to get here. If you disagree with the question, or don't understand it: why?
- JoshuaDavid 3y agoI don't think that's where the difference lies, no. Sampling from a markov model built from English n-gram frequencies will produce non-repeated, mostly grammatical English text, which frequently is even sensible-sounding once n >= 4 or so. But I don't think there's anything going on in language generation beyond "based on the current context, produce an appropriate output token for that context based on the observed and inferred distribution of training inputs". I think that's also how human language generation works, though "the context" for humans includes a lot more than just a few thousand words of text. But I think that the surprising thing about e.g. GPT-4 is how well it does the thing, rather than the fact that it does the thing at all.