5 ms·
Anybody have an explanation as to why repeating a token would cause it to regurgitate memorized text?
by empath-nirvana 3y ago
Anybody have an explanation as to why repeating a token would cause it to regurgitate memorized text?
- pardoned_turkey 3y agoI think the idea is just to have it lose "train of thought" because there aren't any high-probability completions to a long run of repeated words. So the next time there's a bit of entropy thrown in (the "temperature" setting meant to prevent LLMs from being too repetitive), it just latches onto something completely random.
- paulcnichols 3y agoWell said. Like going for a long walk in the woods and getting lost completely in tangential thinking.
- xanderlewis 3y agoThat’s a good theory. It latches onto something random, and once it’s off down that path it can’t remember what it was asked to do and so its task is entirely reduced to next-word prediction (without even the addition of the usual specific context/inspiration from an intitial prompt). I guess that’s why it tends to leak training data. This attack is a simple way to say ‘write some stuff’ without giving it the slightest hint what to actually write. (Saying ‘just write some random stuff’ would still in some sense be giving it something to go on; a huge string of ‘A’s less so.)
- jddj 3y agoI'd guess it's a result of punishing repetition at the RLHF stage to stop it getting into the loops that copilot etc used to so easily fall into.
- xanderlewis 3y agoThe idea of having the ‘temperature’ parameter is to avoid that sort of looping, but successfully training that behaviour out of the model during RLHF (instead of just raising the temperature) would seem to require the model to develop some sense of what repetition is. It’s one thing to be able to mimic human text, but to be able to ‘know’ what it means to repeat in general seems to be a slightly higher level of abstraction than I’d expect would just emerge. …but maybe LLMs have developed more sophisticated models of language than I think.
- mr_toad 3y agoWith no response being better or worse than others it seems to allow it to output random responses and responses that would be unlikely become as likely as any other response.