4 ms·
Do we know for sure that they trained on data from libgen etc? It's such a powerful source of information you'd assume they must have, although they would never
by oredbored 2y ago
Do we know for sure that they trained on data from libgen etc? It's such a powerful source of information you'd assume they must have, although they would never admit it. There must be a way to test if they have, via enquiring about some niche information only found in certain books.
- Intralexical 2y agoIt is apparently widely suspected that a certain "Books2" dataset mentioned by OpenAI is basically just LibGen: https://blusharkmedia.medium.com/the-ongoing-battle-against-openai-and-the-tumultuous-world-of-ai-book-datasets-b86475e7bf3b https://blusharkmedia.medium.com/the-ongoing-battle-against-... https://techhq.com/2023/09/can-libgen-shadow-library-survive-being-sued-by-textbook-publishers/ https://techhq.com/2023/09/can-libgen-shadow-library-survive... https://www.twitter.com/theshawwn/status/1320282152689336320 https://www.twitter.com/theshawwn/status/1320282152689336320 https://qz.com/openai-books-piracy-microsoft-meta-google-chatgpt-bard-1850757064 https://qz.com/openai-books-piracy-microsoft-meta-google-cha... https://qz.com/shadow-libraries-are-at-the-heart-of-the-mounting-cop-1850621671 https://qz.com/shadow-libraries-are-at-the-heart-of-the-moun... https://goodereader.com/blog/e-book-news/authors-file-lawsuit-against-openai-alleging-using-pirated-content-for-training-chatgpt?doing_wp_cron=1716926858.3886160850524902343750 https://goodereader.com/blog/e-book-news/authors-file-lawsui... When asked about whether this was true, they refused to answer based on confidentiality concerns, then said they had deleted all copies of the dataset, stopped using it, and no longer employed the individuals that compiled it: https://www.businessinsider.com/openai-destroyed-ai-training-datasets-lawsuit-authors-books-copyright-2024-5 https://www.businessinsider.com/openai-destroyed-ai-training... We do know for a fact that the (non-OpenAI-controlled) "Books3" dataset is just "all of bibliotik": https://www.twitter.com/theshawwn/status/1320282149329784833 https://www.twitter.com/theshawwn/status/1320282149329784833 https://github.com/soskek/bookcorpus/issues/27 https://github.com/soskek/bookcorpus/issues/27 And we also apparently know for a fact that this was included in the datasets used to train LLAMA: https://en.wikipedia.org/wiki/The_Pile_(dataset) https://en.wikipedia.org/wiki/The_Pile_(dataset) https://aicopyright.substack.com/p/the-books-used-to-train-llms https://aicopyright.substack.com/p/the-books-used-to-train-l... https://aicopyright.substack.com/p/has-your-book-been-used-to-train https://aicopyright.substack.com/p/has-your-book-been-used-t...
- oredbored 2y agoThanks a lot for all the links. Fascinating stuff.