5 ms·
Something I've noticed that both GPT-2 and GPT-3 tend to do is get stuck in a loop, repeating the same thing over and over again. As if the system was relying o
by thamer 5y ago
Something I've noticed that both GPT-2 and GPT-3 tend to do is get stuck in a loop, repeating the same thing over and over again. As if the system was relying on recent text/concepts to go to the next utterance, only getting into a state where the next sentence or block of code being produced is one that has already been generated. It's not exactly uncommon.
What causes this? I'm curious to know what triggers this behavior.
Here's an example of GPT-2 posting on Reddit, getting stuck on "below minimum wage" or equivalent: https://reddit.com/r/SubSimulatorGPT2/comments/engt9v/my_former_boss_has_been_paying_me_less_than/ https://reddit.com/r/SubSimulatorGPT2/comments/engt9v/my_for...
(edit) another example from the GPT-2 subreddit: https://reddit.com/r/SubSimulatorGPT2/comments/en1sy0/im_going_through_a_period_of_questioning_identity/fdt4arh/ https://reddit.com/r/SubSimulatorGPT2/comments/en1sy0/im_goi...
With GPT-3, I saw GitHub Copilot generate the same line or block of code over and over a couple of times.
- not2b 5y agoLimited memory, as the article points out. It doesn't remember what it said beyond a certain point. It's a bit like the lead character in the film "Memento". A very long time ago (early 1990s) I wrote a much simpler text generator: it digested Usenet postings and built a Markov chain model based on the previous two tokens. It produced reasonable sentences but would go into loops. Same issue at a smaller scale.
- Abrownn 5y agoThis is exactly why we stopped using it. Even after fine tuning the parameters and picking VERY good input text, it still got stuck in loops or repeated itself too much even after 2 or 3 tries. It's neat as-is, but not useful for us. Maybe GPT-4 will fix the "looping" issue.
- xwolfi 5y agoHave you talked to an angry redneck in a dank pub ? He'll do the same. And not just him, but whomever is expected to speak but has no high value things to add, and is a little bit dumb/drunk to count how many times he said something and stop at 3. Maybe information, to be interesting to us, has to be novel, while GPT-3 may not model for listener's interest (like you when you re drunk) and only produces the best it can express in a given input context ? And sometimes, maybe repeating 34 times the same thing is good if no new input changes the fundamentals, just not very interesting for a signal dampener like our brain who starts losing focus when novelty disappears from the signal? It s like imagine a political debate around building a bridge between a truck driver who wants to go faster and a bird watcher who wants birds to keep their habitat close to his home. There's no input that can change the fundamentals and it would be expected that after a few loops, no brain could find anything to add and just repeat forever the same thing: but the birds must be close to me or I lose my life's meaning, but the bridge must be built there or I cant optimize my route. The only thing we do is put a time stop and say "ok we got it, now everyone in the public can map their own constraint to the discussion and vote".
- skybrian 5y agoNo, that's not it at all. GPT-3 loops within a sentence, creating an infinitely long sentence of gibberish.
- trevyn 5y agoI’ve found the OpenAI repetition penalty settings to work quite well. Also, prompt writing is an art.
- d13 5y agoHere’s why: https://www.gwern.net/GPT-3#repetitiondivergence-sampling https://www.gwern.net/GPT-3#repetitiondivergence-sampling
- layer8 5y agoThis seems to say that nobody really understands why.
- andreyk 5y agoThis is a common problem with language models in general, and the reason is not that well understood. This paper A Theoretical Analysis of the Repetition Problem in Text Generation (https://arxiv.org/abs/2012.14660 https://arxiv.org/abs/2012.14660) seems to offer a principled answer. Basically the probability maximizing search procedure for text generation can get stuck in loops where the most likely next statement is similar or same to before. I'm no NLP researcher so I don't have easy intuition on it, but that paper seems like a good read.
- thamer 5y agoThanks for this paper! This is exactly what I was looking for.
- sjg007 5y agoI've found that autoencoder and VAE predictions tend to converge. This might be a similar thing. If anyone has advice on preventing that let me know.. All I can think of is add some random noise into the input or making it more GAN like... hey that's an idea.