4 ms·
They are literally not allowed to publish those scans. Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright
by ronsor 2mo ago
They are literally not allowed to publish those scans.
Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.
- fabian2k 2mo agoThe books get trashed because loose pages are much easier to scan.
- beej71 2mo agoDepends on the book. I'd wager the majority of the rare books are in the public domain. But they won't share them because they don't want their competitors to have the data.
- UtopiaPunk 2mo agoIf they are in the public domain, they could publish them. How many books are in the public domain but have never been digitized and shared in a public archive? And if the books are not in the public domain, then they should not be allowed to train their AI models with the material without some kind of license or agreement with the owner of the copyright.
- deleted 2mo ago[deleted]
- ethbr1 2mo ago> If they are in the public domain, they could publish them. That'd be an easy fix with new law -- if you are an AI company with book data, you have a burden to make openly available (or require your suppliers to) all public domain book scans.
- toomuchtodo 2mo ago> They are literally not allowed to publish those scans. Emphasis mine. Just provide a digital copy, no questions asked. They will distributed to various archives globally. I'll pay for the drives and shipping. I understand and can appreciate the potential liability, and am willing to launder it to preserve the subject collection(s) and dataset(s).