3 ms·
Google Books already does exactly this. It has a library of the full text of millions of books. Users can search for a passage of text and google will display t
by subroutine 3y ago
Google Books already does exactly this. It has a library of the full text of millions of books. Users can search for a passage of text and google will display the paragraph where the passage is found.
https://books.google.com https://books.google.com
example:
https://i.ibb.co/DCxJpHN/IMG-3143.jpg https://i.ibb.co/DCxJpHN/IMG-3143.jpg
- palata 3y agoGoogle does not provide the full book, does it? Exactly like they could provide a few seconds of a song, but not the song in its entirety.
- subroutine 3y agoNo, they don't provide the full book, just a few sentences before and after your search prompt (same as Prosecraft). In both cases, however, if you had the patience, you could search the last few words of the text returned from your prior query and slowly work your way through the entire book.
- harshreality 3y agoA few sentences? For most books I've seen, it's a few pages. Google will block you from retrieving more pages from the same book eventually. Using a VPN and a different account may get around one limit, but I experimented with multiple VPNs and browsers once, and although I was able to get a majority of a book's pages, after that google stopped showing me full previews of any of the remaining pages no matter where the request came from.
- subroutine 3y agoIt shows you a few pages if you are previewing the book (i.e. "look inside"). But if you are using search, it will show you where your search query shows up in the book, no matter what page the search query is found. This means you could theoretically search a book sentence by sentence, and it will eventually have shown you the entire book. I'm not claiming this is an efficient or practical way to game the system and read books, only that google books does contain the full copy of the book text and can reveal the contents of any passage. This is basically how Prosecraft works (at least what i glean from the article) - it doesn't let you read a whole book, even though it may contain a representation of the full text.
- palata 3y agoSure. I really did not mean that specifically for Prosecraft. But the article questions why authors are attacking Prosecraft "because it does no harm". My answer is that authors don't (and can't, really) make the difference in a per-case basis. At this point what they see is that LLMs trained on their copyrighted material are able to generate similar material thanks to their copyrighted material that was used in the training (that is important!), and they see that they won't get paid for that. Of course they are scared, and they should be. And of course they will now start attacking everything that looks like it is using their copyrighted material as training data. I really don't get why the engineering world does not get this: LLMs have the potential to ruin people's jobs, it is not clear at all that this is legit (IMO LLMs could not do it without the copyrighted material they used for training, therefore they are derivatives of the original work), and those people are rightfully scared.
- subroutine 3y agoWhy didn't you just say that, instead of posing a hypothetical about software that may itself contain full book text which can be used to display (in this case fair-use) passages to end users? lol I think the disconnect between your point of view and mine is that I see "training an LLM on copyrighted text" the same as a person reading copyrighted text, which is perfectly legal. And I see violating copyright as a person or LLM reproducing copyrighted work (illegal). But using other works as inspiration for something novel shouldn't be considered illegal, whether a person or LLM produced the work. I would even be fine with literature being treated more like music, where reproducing the essence of a piece of work (i.e. doesn't have to be a word for word reproduction) is considered a violation. But if the LLM creates something completely new, how is that a derivative work / infringement?
- palata 3y ago> Why didn't you just say that, instead of Because I answered to a post that was talking about drawing the line for fair use. I just shared my view of how I see it. To me, OpenAI should be responsible for not giving copyrighted material to users if they are not allowed to do it. This means that they should be sued every single time someone manages to extract what is considered as copyrighted material from their software. Because the authors never gave them that right. You Google Books example is different: the most obvious difference between that Google Books does not pretend that it is their content: they clearly say "here is a passage of this book". > I see "training an LLM on copyrighted text" the same as a person reading copyrighted text Yes, I think that is the main discussion point around LLMs. My point is that machines are not humans, and therefore they should not be blindly treated like humans. We should think about the consequence of the machines doing what they do, and decide whether that is legal or not in our society. Otherwise we would give machines the right to vote ("humans can vote, I don't see why machines couldn't").