4 ms·
Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of the
by lulzury 2y ago
Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models?
This also begs the question, does OpenAI really honor their privacy controls and not use user information to train their models if the user opts out? It seems most companies are operating in "ask for forgiveness than permission"-mode as they scramble to stay competitive in the AI race. [0]
[0] https://news.ycombinator.com/item?id=40127106 https://news.ycombinator.com/item?id=40127106
- deleted 2y ago[deleted]
- fallingsquirrel 2y ago> Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models? Sure, in much the same way if you save a JPEG at 75% quality the data in the image is still encoded. But if you repeat a lossy encoding over and over without saving the original, well... wikipedia has a nice visualization of what happens: https://en.wikipedia.org/wiki/Generation_loss https://en.wikipedia.org/wiki/Generation_loss
- Legend2440 2y agoThey do not train newer models on output from older versions. In fact, they deliberately try to filter out LLM-generated text from their scraped internet datasets.