6 ms·
To be clear, the issue is not that the books were used to train Claude, but that they were pirated.
by jdlshore 3mo ago
To be clear, the issue is not that the books were used to train Claude, but that they were pirated.
- EmoteSupportBot 3mo agoA critical distinction, because they were going to to find terabytes of not pirated books to train on that contained the sum history of humanities knowledge /s
- amanaplanacanal 3mo agoThey could have purchased the books instead. It was easier to pirate.
- Aurornis 3mo agoThey actually did this. > Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
- modeless 3mo agoGreat, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!
- fluoridation 3mo ago>instead of allowing anyone to train on already scanned books for free That would be pirating. So your complaint is that they didn't do more piracy?
- modeless 3mo agoMy complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors. To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
- jamesjhare 3mo agono we should destroy the works of these ghouls and support humans instead of this destructive and useless technology the people operating frontier labs are bad people they cannot be trusted in any way the best solution to them would be to send them to monster island (even though it's really a peninsula)
- nextaccountic 3mo ago> Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire. Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though) If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
- blackqueeriroh 3mo agoNo, it’s proof purchase of how stupid the publishing industry is. Maybe publishing houses should just pay authors good money, like a goddamn salary, and get a book out of them every few years.
- bandrami 3mo agoThey can't for the same reason that cab companies can't make their drivers employees: they would have to employ far, far fewer of them than they do on contingency.
- antisthenes 3mo agoThe AI craze not only destroyed books, but many small websites who couldn't bear the load of constant scraping, or many communities that took open forums and took them offline or put them behind closed doors. There is less publicly available knowledge now on the Internet than there has been 3 years ago.
- JAlexoid 3mo agoIt reads like you're in favor of banning resale and lending(aka libraries) of books... because the authors aren't compensated. There's such a thing as fair use and digitizing privately owned printed material is absolutely legal... including for corporations.
- BikiniPrince 3mo agoHow is this ruled as piracy then? I am confused.
- kg 3mo agoThey also pirated the books
- jamesjhare 3mo agowell if they made a PDF copy to process they violated copyright
- qq66 3mo agoThey bought, scannned, trained from, and destroyed millions of paper books, which was ruled legal. This lawsuit was for training from LibGen.
- scotty79 3mo agoThis is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that. But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
- gruez 3mo ago>This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that. That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
- sillysaurusx 3mo agoIt's an unfortunate outcome. Now to be a big player in AI, you have to have enough capital to buy your own library worth of books and digitize them. (Fun fact: a pallet of books is called a "gaylord," and they buy hundreds of gaylords.) I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/ https://sites.research.google/trc/about/). Now we're in a world where you have to have dozens of millions in capital to do substantial work. I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
- kryogen1c 3mo ago> Fun fact: a pallet of books is called a "gaylord," A Gaylord is a type of box that fits on a pallet. There are multiple ways to palletize products, like shrink wrapping or metal banding
- spaqin 3mo agoLadder-pulling at its best. It's also easier to swallow the fine once you've launched a successful product after pirating the books.
- usef- 3mo agoJudging by the current comment section, most HN people seem to want the ladder pulled (an observation, not agreement)
- satvikpendem 3mo agoAs this comment states, it may not be the piracy that is the issue, but the keeping of the books forever rather than just for the purpose of training. Regardless, I believe people will do this in a wink wink nudge nudge sort of way anyway, as no one releases their training data, because we all know where it comes from. https://news.ycombinator.com/item?id=48996652#49004015 https://news.ycombinator.com/item?id=48996652#49004015
- jamesjhare 3mo agohow much of your economic output are you comfortable with companies like Anthropic stealing to put you out of work? at least in Player Piano they paid the workers who made the cassette tapes that made the robots work. our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set. they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.
- blackqueeriroh 3mo agoAll of it.
- ori_b 3mo agoIf you believe information deserves to be free, and if most of your earnings were from information that wasn't given away for free -- well, if you want people to give up their ill gotten gains, maybe you can start by setting an example. So, mind sending me your bank account information? I'll promise to make good use of it.
- JAlexoid 3mo agoYou seem to think that AI companies sell you content, which is false. You get a service. The service is using their compute power to run a model and their scientists to build the model.
- ori_b 3mo agoNo. I think AI companies consume the result of intellectual work that should typically be paid for. If someone doesn't believe intellectual work should be paid for so that others can benefit, I invite them to lead the way.
- mountainriver 3mo agoThis is why you should always distill your models from a competitor. Let them take on the liability
- wraptile 3mo agoWith this it's becoming very clear that we're moving past information copyright of today and the only copyright that'll remain will be brand/trademark shaped. This might be a good thing right? Information remains free while people's effort remains protected (assuming fair governance).
- inigyou 3mo agoThat would be trademark not copyright
- nicce 3mo agoWhich is another issue. In the context of pirating, it should be also an issue, because it is a benefit from the crime.