5 ms·
From a quick web search I can find that there are book review sites that allow users to enter and rate verbatim "quotes" from books. This one [1] contains ~2000
by fuzzbazz 1y ago
From a quick web search I can find that there are book review sites that allow users to enter and rate verbatim "quotes" from books. This one [1] contains ~2000 [2] portions of a sentence, a paragraph or several paragraphs of Harry Potter and the Sorcerer's Stone.
Could it be plausible that an LLM had ingested parts of the book via scrapping web pages like this and not the full copyrighted book and get results similar to those of the linked study?
[1] https://www.goodreads.com/work/quotes/4640799-harry-potter-and-the-philosopher-s-stone https://www.goodreads.com/work/quotes/4640799-harry-potter-a...
[2] ~30 portions x 68 pages
- aspenmayer 1y agoSure, why not? lol https://www.reddit.com/r/DataHoarder/comments/1entowq/i_made_a_tool_to_scrape_magazines_from_google/ https://www.reddit.com/r/DataHoarder/comments/1entowq/i_made... https://github.com/shloop/google-book-scraper https://github.com/shloop/google-book-scraper The fact that Meta torrented Books3 and other datasets seems to be by self-admission by Meta employees who performed the work and/or oversaw those who themselves did the work, so that is not really under dispute or ambiguous. https://torrentfreak.com/meta-admits-use-of-pirated-book-dataset-to-train-ai-240111/ https://torrentfreak.com/meta-admits-use-of-pirated-book-dat...
- redox99 1y agoBooks3 was used in Llama1. We don't know if they used it later on.
- aspenmayer 1y agoMy comparison was illustrative and analogous in nature. The copyright cartel is making a fruit of the poisonous tree type of argument. Whatever Meta are doing with LLMs is doing the heavy lifting that parity files used to do back in the Usenet days. I wouldn’t be surprised if BitTorrent or other similar caching and distribution mechanisms incorporate AI/LLMs to recognize an owl on the wire, draw the rest just in time in transit, and just send the diffs, or something like that. The pictures are the same. All roads lead to Rome, so they say.
- aprilthird2021 1y agoAll of the major AI models these days use "clean" datasets stripped of copyrighted material. They also use data from the previous models, so I'm not sure how "clean" it really is
- dragonwriter 1y ago> All of the major AI models these days use "clean" datasets stripped of copyrighted material. Which of the major commercial models discloses its dataset? Or are you just trusting some unfalsifiable self-serving PR characterization?
- aprilthird2021 1y agoIt's from my personal experience in the industry
- aspenmayer 1y agoWhat are your thoughts on the origin of the LLaMA leak? It's interesting that the training data was torrented, and so was the leak. Perhaps we will never know? For the OSINT folks, not a lot to go on, or maybe a lot, depending? https://en.wikipedia.org/wiki/Llama_(language_model)#Leak https://en.wikipedia.org/wiki/Llama_(language_model)#Leak https://archived.moe/g/thread/91848262#p91850335 https://archived.moe/g/thread/91848262#p91850335 https://github.com/meta-llama/llama/pull/73/files https://github.com/meta-llama/llama/pull/73/files
- aprilthird2021 1y agoI don't really know much about that, sorry
- aspenmayer 1y agoI didn’t ask for info, I asked for your views. I gave you all the info anyone has publicly, so you have enough to comment. I suspect that it was a limited hangout self-own by Meta to claim that they aren’t responsible, and then they are doing research on a leaked LLM that they developed, but then was leaked, so they can claim that the subsequent research is not tainted by the fruit of the poisonous tree legal doctrine. Or, their torrent client or other software on the same machine had 0-days and they got hacked by someone on the Books3 swarm or knowledgeable of what IPs were connecting to it. I appreciate your posts and I am replying to you to humbly ask you to post more. :P
- paxys 1y agoMeta has trained on LibGen so we don't really need to speculate. https://www.wired.com/story/new-documents-unredacted-meta-copyright-ai-lawsuit/ https://www.wired.com/story/new-documents-unredacted-meta-co...
- aprilthird2021 1y agoThis is in fact mentioned and addressed in the article. Also, there is pretty clear cut evidence Meta used pirated book data sets knowingly to train the earlier Llama models