3 ms·
Try googling sentences.
by nmca 8y ago
Try googling sentences.
- minimaxir 8y agoHmm, I tested a few sentences and it didn't turn up any exact matches (aside to this article), so maybe I'm wrong. With a temperature of 0.7/1.0, that's enough for sufficiently random text I suppose. (the raw, uncurated generated text using the smaller model is a bit more random: https://raw.githubusercontent.com/openai/gpt-2/master/gpt2-samples.txt https://raw.githubusercontent.com/openai/gpt-2/master/gpt2-s...)
- wuthefwasthat 8y agoThose samples are from the large model (GPT-2)! Regarding memorization vs. generalization, see our paper for more analysis.
- yorwba 8y agoHave you tried to determine which parts of the training data contributed to the model generating a certain output? I wonder whether the model avoids reproducing exact matches of the training data by splicing several similar articles together. For example, the generated text about the Civil War mentions that Thomas Jefferson Randolph [0] was named after his grandfather, the president. But is the wording mostly influenced by articles talking about that specific fact, or does it draw from more general examples of someone being named after their grandfather? [0] https://en.wikipedia.org/wiki/Thomas_Jefferson_Randolph https://en.wikipedia.org/wiki/Thomas_Jefferson_Randolph
- zalo 8y agoDoes it seem like there will be any way to go backwards from the sample to the prompt? From a safety perspective, it would be useful to see what prompt a piece of text might have been generated with...