4 ms·
GPT-3's coherency inevitably degrades when it writes anything of length, because its "working memory" (context window) is only a few hundred words. (2048 BPE to
by orost 6y ago
GPT-3's coherency inevitably degrades when it writes anything of length, because its "working memory" (context window) is only a few hundred words. (2048 BPE tokens, which are characters, clusters of characters, or words)
So by the time it's approaching the middle of a 2k word essay, it's starting to forget the beginning (including any prompt that tells it what it's meant to be writing about), and by the time it's writing the ending, it has forgotten the middle. Except insofar as their contents are reflected in further paragraphs.
Obviously that's a crippling limitation, but certainly a fixable one (by throwing more $$$ at the problem if nothing else). I am very curious to see long-form output of a GPT with a longer attention span.
- Baeocystin 6y agoDo we know what the cost is for increasing the window size? I'd imagine that each additional token of range requires ever more power to properly integrate, but that is only a hunch.
- orost 6y agoI believe it's quadratic, but that still means making it 10x larger is perfectly possible just by brute force with a little motivation.
- Baeocystin 6y agoInteresting. That's less steep than I would have naively guessed. And you're right, if it's that 'easy', the next iteration will be (1) soon and (2) awesome, in both the positive and negative sense of the word.
- nullc 6y ago> Obviously that's a crippling limitation, FWIW, you can work around this some by doing the writing as a conversation like this. I am writing a story about "xyz" the first paragraph is "qwe". I continue my story about "xyz" the next pargraph is
- nullc 6y agoYou can even get it to maintain it's own 'memory': First start off like normal, and after it's written a bit, rewrite it so that it's repeating a summary critical topics between each paragraph. Then let it go, it'll continue carrying through and even sometimes updating its 'memory'. (not updating as much as I'd like: it's extremely efficient at just copying text from recent history). I think it would be interesting to train GPT3 specifically to work this way: First train GPT3. Then, run back over the training data and use GPT3 to generate running summaries: At each paragraph break add some text that says something like "The most important things about the above text is:" and let it complete that prompt). Then use those running summaries to augment the training data with special symbols that occur nowhere in the input marking the self-commentary parts, omitting the prefix you used to get gpt3 to output it, and train a new network (GPT3') on the augmented data. Then you could make an interface that uses GPT3' and hides the self-commentary from the users. As GPT3 writes it will have a persistent memory that can last as long as the document goes on, updated by itself. Effectively it gives it sparse access to the entire history, but the network itself controls the shape of the access. [Plus a nice thing about GPT generated text is that you can store its confidence too, and use that to weigh the training so that you penalize it less for mispredicting stuff it was unsure of.] You wouldn't have to do anything special to teach it to write this commentary because we already write commentary in English and GPT3 already knows how to do it. Maybe a little more engineering would be useful to guarantee that it will write commentary blocks often enough, but it could be as simple as making sampling prefer emitting an internal monologue block with an increasing bias as the last one approaches falling out of the window.