3 ms·
> If they had bought each book themselves would it be fair use? So this is only about the piracy? The earlier ruling covered exactly that question: - Anthropi
by ijk 1y ago
> If they had bought each book themselves would it be fair use? So this is only about the piracy?
The earlier ruling covered exactly that question:
- Anthropic downloaded many books (from LibGen and elsewhere). This piracy is what the current case is about, and is unrelated to the training.
- Separately, Anthropic bought and scanned a million used books. They trained the AI on this data. This was ruled as fair use, and is not involved in the current case.
- Ajedi32 1y agoThat's very interesting, because it totally makes sense legally, but the practical effect is ludicrously stupid. The law is effectively forcing companies to spend millions re-scanning the same books over and over for no reason. It'd be like if we had a law which stated "Before you can train an AI, you must light 1 million dollars on fire. After that you can do whatever you want.". It serves no purpose but to waste societal resources on nothing.
- JohnFen 1y agoThe cost might reduce the number of entities who can afford to do it, though, which would reduce the amount of abuse.
- aswegs8 1y agoWhy would they need to "re-scan the same books over and over"? It's as simple as they can use the books to train their AI if they bought them.
- Ajedi32 1y agoBecause company A needs to scan the books, then company B wants to train their AI so they need to scan the same books, then company C wants to train their AI so they need to scan the same books... etc. It would be one thing if they were buying "used" digital copies of the books, but the fact that this is only legal with scanned physical copies makes it extremely wasteful.
- ijk 1y agoIt would probably be legal with digital copies; it's just that book publishers have been very zealous in preventing the existence of a market for used digital books. Copyright has been very silly in the digital realm from the beginning and is unlikely to get less unhinged from reality absent a major overhaul that makes it completely unrecognizable.
- triceratops 1y agoDigital media, in particular its resale, is the one good use case for blockchains that no one seems to be interested in (and don't provide me a link of some obscure project working on it; what can I buy with it?). Probably because it's useful for consumers but not for making money.
- alias_neo 1y agoI don't have any sympathy for big orgs, they can follow the same rules as the rest of us, and should be slapped even harder for this than an individual accused of the same thing, however, I'm curious, why can't they buy digital copies in the first place? Is there some nuance to the law that allows them to scan/copy them if they're physical but not if they're digital?
- ACCount36 1y agoNot every book is readily available as a digital copy. Things like textbooks, older technical books or just books that weren't too popular can be easier to source as physical books and scan destructively. A lot of digital copies are also DRM'd to shit - to obtain raw text usable for AI training, you'd have to break DRM. Which isn't that hard, on a technical level - but DMCA exists. DMCA is a shit law that should have been dismantled two decades ago - but as long as it's around, bypassing DRM on things you own can be illegal. Scanning sidesteps that.
- alias_neo 1y agoI'm anti-DRM personally, but I suppose in this case we could argue it's serving its purpose, it's just that workarounds have been found in the form of scanning physical books. If no physical copies existed and there were only DRMd digital copies of everything, the companies scanning books for AI training would be forced to work out some deal with the DRM-overlords to have it removed for their use. That (I think) would be a net benefit as hopefully the authors would get paid too.
- bostonsre 1y agoIt needs to hook into the existing legal book supply chain so that authors could potentially get compensated (I doubt they do for used book resale tho..).
- Eric_WVGG 1y ago"no reason"? Try telling that to the people who wrote the books.
- tpmoney 1y agoI’ve been thinking recently that an overhaul to the copyright system could solve this. Return to a very low default (10 years? 20?). Allow extensions but a requirement for extension is submitting the work to a government managed digital data set that is licensed out to people to use as training data for these sorts of systems (or anything else a massive digitized cataloged library could be useful for). Licensing is some nominal amount of money and the revenue from that is distributed to copyright holders who have submitted their works in proportion to the recency and volume of content (with some cap to avoid flooding the system with content just to get more payouts. I’m sure there’s lots of unintended problems with this, but it does feel like a common base set of training data like this is exactly the sort of thing the government can and should do.
- alias_neo 1y ago> The law is effectively forcing companies to spend millions re-scanning the same books over and over for no reason Would anyone agree if you replaced companies with people in that argument? Why shouldn't a company follow the same rules as everyone else just because the scale at which they're doing it is so large? I'd argue a company doing something like this should be forced to buy the books NEW and benefit the authors, and if they're found guilty of copyright infringement they should be punished at a scale a few orders of magnitude larger than an individual would be. > Before you can train an AI, you must light 1 million dollars on fire If I want to train an AI, I probably need to spend a larger part of my budget as an individual to do so than an org, should I be given the resources for free or severely discounted because I want to make money out of it? I suppose one _could_ argue in favour of such a practice if it was going to benefit society as a whole, but is it?
- Ajedi32 1y agoI'm not saying companies should follow different rules than people, I'm saying the rules as written make no sense. This particular example just happens to make that fact more readily apparent due to the sheer scale of the needless waste involved.
- alias_neo 1y agoI'm anti-DRM myself, but someone else could argue that the rules are partly doing their job; preventing companies from just gobbling up digital copies, it just happens that they have he resources to take advantage of a loophole by scanning the books in themselves. The best solution I can come up with would be a digital library where one org, say the internet archive has scanned everything once, then they're charge a licence fee to these orgs to ingest a copy, and the part of the payment goes to the author, no big wastage, the information gets archived and the orgs pay their share.
- watwut 1y ago> Before you can train an AI, you must light 1 million dollars on fire. I mean, demanding you pay money to the source of data in your quest to create a monopoly you are pretty much guaranteed to abuse later on while becoming filthy rich is not exactly unfair.
- troyvit 1y ago> The law is effectively forcing companies to spend millions re-scanning the same books over and over for no reason. Oh but the reason is that they're now making $3 billion/year, partially because of those books. I see an argument for the inefficiency behind having to rescan books that are already scanned, but not the cost. If there was a way to buy pre-scanned books from Google Books or whatever then I somewhat see where you're coming from. I argue that there were positive effects of Anthropic having to buy and scan physical books: * The choices people made choosing which physical books to buy and scan helped make Claude what it is. Personally I sense a difference between Claude and OpenAI and Gemini, and part of it comes down to the choices they made in training material. Sorry to go on and on, but how many choices here were made because it was a rainy day and the trains were down, so an intern went to bookstore A instead of bookstore B? * While buying the books used didn't help the authors it helped the struggling bookstores selling their books. Literal dollars into the hands of local workers. When I fast forward to today and see how LLM companies are literally stealing the energy from the communities their data centers are based in, and polluting them with shitty power plants I can at least think of that as one positive outcome, even if it only happened once. As far as the 7 million+ books Anthropic didn't pay for, their series B in 2022 brought in $580 million. They could have afforded those books.
- tedivm 1y agoThe law isn't forcing people to do this, economics are. Nothing about the law forces people to use physical books, just that they actually pay for the books instead of stealing them. The company thinks they can get away with this cheaper than negotiating for a digital copy of the book, so that's what they are doing.
- BobaFloutist 1y agoI mean I don't think the law precludes them legally purchasing ebooks.