4 ms·
I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and o
by thread_id 1mo ago
I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.
https://en.wikipedia.org/wiki/Google_Books https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-rules-that-google-book-scanning-is-fair-use/ https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-millions-of-print-books-to-build-its-ai-models/ https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
- titzer 1mo agoLike all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space. It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.
- nazgulsenpai 1mo agoThey are also an AI company now. Why would they stop?
- al_borland 1mo agoThey could use them as training data, without providing access the actual books.
- jimmaswell 1mo agoAt least we would all benefit from the books this way, so long as legal nonsense keeps the scans unavailable to the public.
- al_borland 1mo agoI don’t think locking the content of rare books away in the hands of corporations who only give us access to tools trained on the books, and not the actual book, is a good path forward. This doesn’t incentivize them to be good stewards of this data and making anything in the public domain available. It incentivizes less access to the source material, having to blindly trust their tools, and is effectively automating plagiarism.
- JKCalhoun 1mo agoAlso, why would they ever share their collection? Book scans, secreted away, are worthless to the public.
- greyw 1mo agoIt's data for their AI pipeline. Basically digital gold.
- ktm5j 1mo agoI think you're missing the point. Maybe I'm wrong, but I'm pretty sure they're trying to point out that this book scanning doesn't need to be destructive. I'm not sure if AI companies are using a scanning method that damages the book or not, but they destroy the books after scanning to avoid copyright issues (ie they aren't duplicating the books). This Google project seems to demonstrate that this isn't actually necessary, and that the AI companies are just doing it out of laziness.
- shagie 1mo agoIf it isn't scanned destructively, is it a liability? Can you do anything else with the book? What are its costs for storage in a way that retains the value of the book? If the assets of the warehouse are sold to another company (see also https://paizo.com/blog/paizo-restructuring-a-difficult-update-about-our-future https://paizo.com/blog/paizo-restructuring-a-difficult-updat... ), what are your obligations for the format shifted copy that you retain? These questions imply that there's a liability that exists when retaining the original that has little value to the company. And they (the books) aren't assets that can be resold. It's easier (and cheaper), has no ongoing costs for physical storage, and answers those questions without creating legal entanglements for the future company.
- ktm5j 1mo agoQuestions don't imply anything, you're just asking things. I don't have the answers to your questions, but the point is that Google was able to scan books without destroying them while still avoiding legal consequences (with some effort).
- shagie 1mo agoThe physical books that were scanned by Google were returned to the libraries. https://btaa.org/library/programs-and-services/book-search/faq https://btaa.org/library/programs-and-services/book-search/f... > Will scanning harm the books? > No. Google developed innovative technology to scan the content without harming the books. Any book deemed too fragile will not be scanned by Google, but may be treated by expert library staff. Once scanned, all print volumes are returned to the library collections. That was an inherently different goal (borrow the books from the library, scan them, and return them) than the Bartz v. Antrophic ruling. https://cases.justia.com/federal/district-courts/california/candce/3%3A2024cv05417/434709/231/0.pdf https://cases.justia.com/federal/district-courts/california/... > Storage and searchability are not creative properties of the copyrighted work itself but physical properties of the frame around the work or informational properties about the work. See Texaco, 802 F. Supp. at 14 (physical), aff’d, 60 F.3d at 919; Google, 804 F.3d at 225 (informational); Sony Corp. of Am. v. Universal City Studios, Inc. (“Sony Betamax”), 464 U.S. 417, 447 (1984) (rightful interests). In Texaco, the court reasoned that if a purchased scientific journal article had been copied “onto microfilm to conserve space, this might [have been] a persuasive transformative use.” 802 F. Supp. at 14 (Judge Pierre Leval), aff’d, 60 F.3d at 919 (reducing “bulk[ ]” “might suffice to tilt the first fair use factor in favor of Texaco if these purposes were dominant“). In Google Books, the court reasoned that a print-to-digital change to expose information about the work was transformative. Google, 804 F.3d at 225 (Judge Pierre Leval). And, in Sony Betamax, the Supreme Court held that making a recording of a television show in order to instead watch it at a later time was copying but did not usurp any rightful interest of the copyright owner. 464 U.S. at 447, 455. Important to the Supreme Court’s reasoning was the expectation that most such copiers would not distribute the permanent copies of the work. Finally, in A&M Records, Inc. v. Napster, Inc., our court of appeals recognized the reasoning just explained, and therefore rejected by contrast a digitization effort that was touted as space-shifting but in fact resulted in the multiplication of copies shared with outsiders through a file-sharing service. 239 F.3d 1004, 1019 (9th Cir. 2001), aff’g in this part 114 F. Supp. 2d 896, 912–13, 915–16 (N.D. Cal. 2000) (Judge Marilyn Hall Patel) (citing Sony Betamax and Texaco). > Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others). --- The AI training isn't borrowing from libraries and returning from libraries. Instead, it is buying a book, format shifting, and retaining that format shifted version from its own use. The company can't do anything else with the book once they've format shifted it. They can't donate it and they can't resell it. In that light, destructively scanning the book is the best option. There is no value in trying to non-destructively scan it because otherwise all it would do is sit in a warehouse and cost money to pay for storage of something they can't sell.
- dblohm7 1mo ago> It's great they did this, but the Google that is today cannot be trusted with data of public value anymore. They never should have been trusted in the first place.
- probably_wrong 1mo agoWhenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included. I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.
- p0w3n3d 1mo agoQuod licet Iovi, non licet bovi Big companies will read up the books and make their AI recite them from memory, but Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)
- misnome 1mo ago> Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent) No, this is what they were doing before, but they explicitly started lending out "unlimited" copies, which is why they got sued.
- ndiddy 1mo agoThat's why they got sued, but the suit is mainly over whether controlled digital lending is legal at all rather than their "emergency library". Archive.org lost the case on summary judgment, meaning that they could not come up with a single fair use argument for CDL that the judge found compelling enough to let the case go to trial. The full judgment is here https://storage.courtlistener.com/recap/gov.uscourts.nysd.537900/gov.uscourts.nysd.537900.188.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.nysd.53... but here's a couple excerpts: > The crux of IA's first factor argument is that an organization has the right under fair use to make whatever copies of its print books are necessary to facilitate digital lending of that book, so long as only one patron at a time can borrow the book for each copy that has been bought and paid for. See Oral Arg. Tr. 31:10-15. But there is no such right, which risks eviscerating the rights of authors and publishers to profit from the creation and dissemination of derivatives of their protected works. See 17 U.S.C. §§ 106(1), (2). IA's wholesale copying and unauthorized lending of digital copies of the Publishers' print books does not transform the use of the books, and IA profits from exploiting the copyrighted material without paying the customary price. The first fair use factor strongly favors the Publishers. > In this case, there is a "thriving ebook licensing market for libraries" in which the Publishers earn a fee whenever a library obtains one of their licensed ebooks from an aggregator like OverDrive. Pls.' 56.1 ¶¶ 577-578. This market generates at least tens of millions of dollars a year for the Publishers. Id. ¶¶ 170, 172. And IA supplants the Publishers' place in this market. IA offers users complete ebook editions of the Works in Suit without IA's having paid the Publishers a fee to license those ebooks, and it gives libraries an alternative to buying ebook licenses from the Publishers. Indeed, IA pitches the Open Libraries project to libraries in part as a way to help libraries avoid paying for licenses. See Pls.' 56.1 ¶ 383 (presentation IA gave to libraries asserting that pairing with IA means that "You Don't Have to Buy It Again!"); id. ¶ 382 (different presentation promising that the Open Libraries project "ensures that a library will not have to buy the same content over and over, simply because of a change in format"). IA thus "brings to the marketplace a competing substitute" for library ebook editions of the Works in Suit, "usurp[ing] a market that properly belongs to the copyright-holder."
- toomuchtodo 1mo agoInternet Archive version: https://openlibrary.org/ https://openlibrary.org/ Info on where to send books not yet in their collection: https://help.archive.org/help/how-do-i-make-a-physical-donation-to-the-internet-archive/ https://help.archive.org/help/how-do-i-make-a-physical-donat... Mobile apps to determine if they need a book: https://help.archive.org/help/donate-books-app-for-ios-and-android/ https://help.archive.org/help/donate-books-app-for-ios-and-a... Web app: https://archive.org/want/?mode=donation_book https://archive.org/want/?mode=donation_book For example, I donated a copy of Systems Bible (out of print, hard to find imho) and paid for it to jump the digitization queue (https://archive.org/details/systemsbiblebegi0000gall/ https://archive.org/details/systemsbiblebegi0000gall/). The original book will remain stored as a physical backup. It's not fully publicly available of course due to copyright (it will eventually be made public by the Internet Archive once its copyright expires ~2084 and it enters the public domain), which is where shadow libraries|archives like Anna's Archive and Z-Library fill the gap. If you have rare books you would like digitized, archived, and distributed, I am very interested in providing assistance.
- shagie 1mo agoAs an aside... https://www.google.com/books/edition/The_Systems_Bible/mrOsbi2pjXcC?hl=en&gbpv=1&printsec=frontcover https://www.google.com/books/edition/The_Systems_Bible/mrOsb... (and I haven't hit any "you can't read this" limits yet). While the hard copy is a bit pricy for my shelf of curious books, it's also available on kindle. https://www.amazon.com/SYSTEMANTICS-SYSTEMS-BIBLE-John-Gall-ebook/dp/B00AK1BIDM/ref=sr_1_1 https://www.amazon.com/SYSTEMANTICS-SYSTEMS-BIBLE-John-Gall-...
- deleted 1mo ago[deleted]
- jacekm 1mo ago> The project was met with significant legal challenges from authors and publishers which was eventually overcome I don't think they were overcome. As far as I remember Google couldn't make the books available so they abandoned the project. They possess the scans (if they didn't delete them) but they won't be made public.
- chungusamongus 1mo agoIIRC the courts ruled that because it would be implausibe for a person to use google books previews to read an entire work (you'd have to make a whole bunch of separate accounts to do so), it could not plausibly impact the market for that work.
- doctorpangloss 1mo agoThe consensus among copyright lawyers is Google lost Authors Guild v Google, but for some reason the media and Wikipedia do not clearly report it that way.
- wrathofquan 1mo agoI've worked in academic libraries since 2010. I've always felt like it was a mistake for our digital library leadership to put so much trust in Google Books despite their promises to maintain the integrity of libraries (this mostly happened by the way). At the time it was obvious and innovative but over time it was clear Google was establishing a technical precedent to corrode what libraries have the power to do. I'm hoping we can continue to do the good work but it's exhausting.
- b112 1mo agoIf you read the judgement against (I think) openai, the judge said that it was OK to scan the books for LLMs, if they were destroyed afterwards. That is, only one copy of the data existed.
- spandrew 1mo agoAmazon also did this for their "Look Inside" feature. To do this they had to spin up massive digital infrastructure——then realized that they could sell that infrastructure and make more money than from the books they were scanning. This is was what ended up spinning up AWS as a business. I dislike the idea of destroying rare books--but how rare? A digital copy has a lot more benefit.
- janpeuker 1mo agoThe irony of countries blocking Anna's Archive (UK, Italy, Netherlands etc) but then it's the hackers who conserve and steward the books. Because governments can't stop AI companies from shredding history like they did with Google Books _because_ they actually tried to do it _by the book_.
- economistbob 1mo ago[dead]
- etdznots 1mo agoSounds like a great way for the books to be “preserved” in the hands of someone that will never make them accessible, no thank you Google!