8 ms·
It would be incredible for LLMs. Searching it, using it as training data, etc. Would probably have to be done in Russia or some other country that doesn't respe
by namlem 1y ago
It would be incredible for LLMs. Searching it, using it as training data, etc. Would probably have to be done in Russia or some other country that doesn't respect international copyright though.
- sam_lowry_ 1y agoLLMs already use it, dude )
- exe34 1y agoI think one use would be to search for information directly from a book, rather than get a garbled/half-hallucinated version of it.
- jdironman 1y agoYou don't need AI for that. I get the optimistic spirit of what you mean though.
- mdp2021 1y agoOptimized information retrieval of complex text is AI.
- echollama 1y agogarbled/half-hallucinated is probably what you would've gotten 8-12mo ago but now adays im sure with good prompting you can pull value from any book.
- jxjnskkzxxhx 1y agoDo you have a reason to believe this ain't already being done? I would assume that the big guys like openai are already training on basically all text in existence.
- IlikeKitties 1y agoIn fact, facebook torrented annas archive and got busted for it, because of course they did: https://torrentfreak.com/meta-torrented-over-81-tb-of-data-through-annas-archive-despite-few-seeders-250206/ https://torrentfreak.com/meta-torrented-over-81-tb-of-data-t...
- HDThoreaun 1y agoEvery LLM maker probably did the same. Facebook just has disgruntled employees who leaked it
- gpm 1y agoGoogle goes around legally scanning every book they can get their hands on with books.google.com. Legally scanning every paper they can get their hands on with scholar.google.com. I doubt they'd resort to piracy for what is basically the same information as what they've already legally acquired...
- lcnPylGDnU4H9OF 1y agoThat is a good reason to think they did not but it doesn't necessarily override reasons for them to do so. Perhaps it's dubious that the subset of data they could not legally get their hands on is an advantage for training but I really don't know, and maybe nobody does. Given that, Google's execs may have been in favor of similar operations as Facebook's and their lawyers may have been willing to approve them with similar justifications.
- sneak 1y agoDownloading a torrent isn't piracy if you are a license holder for the information that you are downloading.
- gpm 1y ago*If the license you have authorizes you to make a copy in that fashion. But here, Google isn't a license holder. Google doesn't license the text in Google Books (unless something has changed since the lawsuits). Google simply legally acquires (buys, borrows, etc) a copy of the book and does things with it that the US courts have found are fair use and require no license. Incidentally I believe the French courts disagreed and fined them half a million dollars or so and ordered them to stop in France.
- executesorder66 1y ago> or some other country that doesn't respect international copyright though. Like the US? OpenAI et al. don't give a shit.
- TeMPOraL 1y agoThere's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.
- corgi912 1y agoThere's this famous phrase in Russian that was born out of a short interview with a woman, a strong Putin supporter, that's often been used as a sarcastic remark for pointing out someone's double standards and/or hypocrisy. It can be roughly translated to "you don't understand, it's a completely different situation". That's what's constantly on my mind when I'm reading discussions like this one. Everybody and their dog torrenting petabytes of data and getting away with it (Meta is the only one that got caught and they've still gotten away with doing it)? The very same data poor American students were forced to commit suicide over? The same data that average American housewives were sued over for millions of dollars of "damages"? The same data that often gets random German plumbers or steelworkers to pay thousands of euros of "fines" to the copyright mafia so they won't get sued and have their lives ruined? Yet when giant corporations are doing the exact same thing on a massive scale, it's fine? It's not even the same thing, an American student torrenting books isn't making any money off it, while Meta very much is. Of course it's not the same, a simple-minded and poorly educated person like me isn't capable of understanding the difference. You keep believing in your moral superiority, the rest of the world has finally woken up.
- Exoristos 1y agoThere are those who are in charge and those who aren't.
- andrepd 1y ago> Would probably have to be done in Russia or some other country that doesn't respect international copyright though. Incredible, several years of major American AI companies showing that flaunting copyright only matters if it's college kids torrenting shows or enthusiasts archiving bootlegs on whatcd, but if it's big corpos doing it it's necessary for innovation. Yet some people still believe "it would have to be done in evil Russia".
- deleted 1y ago[deleted]
- DataDaoDe 1y agoOP does have an exaggerated statement - its not like there aren't laws in Russia or something and I largely agree with your sentiment. I think there are levels to this though and its pretty clear that Russia is much riskier than the USA when it comes to IP - just look up anything to do with insuring IP risk in Russia (here's one such example: https://baa.no/en/articles/i-have-ip-in-russia-is-my-ip-at-risk https://baa.no/en/articles/i-have-ip-in-russia-is-my-ip-at-r...) Also according to the office of US trade representative, Russia is on the priority watch list of countries that do not respect IP [1] and post 2022, largely due to the war, Russia implemented measures negatively effecting IP rights. [2,3] If you think it isn't the case and Russia is just as risky as the US when it comes to copyright and IP, I would be interested to know why. 1. https://ustr.gov/about/policy-offices/press-office/press-releases/2025/april/ustr-releases-2025-special-301-report-intellectual-property-protection-and-enforcement#:~:text=Eight%20countries%20are%20on%20the,engagement%20during%20the%20coming%20year. https://ustr.gov/about/policy-offices/press-office/press-rel... 2. https://www.papula-nevinpat.com/executive-summary-the-ip-situation-in-russia-and-ukraine/ https://www.papula-nevinpat.com/executive-summary-the-ip-sit... 3. https://www.taftlaw.com/news-events/law-bulletins/russia-issues-decree-affecting-ip-rights-for-unfriendly-countries/?utm_source=chatgpt.com https://www.taftlaw.com/news-events/law-bulletins/russia-iss...
- mdp2021 1y ago> evil In this case and context, a label like "evil" is a twisted interpretation.