6 ms·
I think the headline is a bit misleading. Mets did pirate the works but may be entitled to use them under fair use. It seems like the authors are setting up fo
by TimPC 1y ago
I think the headline is a bit misleading. Mets did pirate the works but may be entitled to use them under fair use. It seems like the authors are setting up for failure by making the case about whether the AI generation hinders the market for books. AI book writing is such a tiny segment what these models do that if needed Meta would simply introduce guard rails to prevent copying the style of an author and continue to ingest the books. I also don’t think AI generated fiction is anywhere near high quality enough to substantially reduce the market for the original author.
- stego-tech 1y agoThe problem is that "harm" as defined by copyright law is strictly limited to loss of sales due to breach of that copyright; it makes no allowment (that I know of) to livelihoods lost by the theft of the work indefinitely, as AI boosters suggest their tools can do (replace people). The way this court case is going, it's an uphill battle for the plaintiffs to prove concrete harm in that very narrow context, when the real harm is the potential elimination of their future livelihoods through theft, rather than immediately tangible harms. As a (creative) friend of mine flatly said, they refuse to use an LLM until it can prove where it learned something from/cite its original source. Artists and creatives can cite their inspirational sources, while LLMs cannot (because their developers don't care about credit, only output) by design. To them, that's the line in the sand, and I think that's a reasonable one given that not a single creative in my circles has been cut payment from these multi-billion-dollar AI companies for the unauthorized use of their works in training these models.
- tedivm 1y agoGithub does give free copilot access to open source developers it considers important enough (which is a pretty low bar). While not the same as actually paying, it's the only example I can think of where the company that used people's copyrighted material actually gave something back to those people.
- _aavaa_ 1y agoWhat you’re describing is the Extend phase of Microsoft’s plan.
- mistrial9 1y agothe educated and erudite can wait in line near the Castle; every day due to the grace of our masters, unused bread from the master's table is available without prejudice. These people can have a fine life, and the people are fulfilled.
- bgwalter 1y agoThey want to track and utilize the new code that those developers are writing. And they want to keep them on GitHub. And they want to claim in potential lawsuits: "See, those developers themselves have used CoPilot, so they approve the copyright infringement."
- subscribed 1y agoYour friend might want to check Perplexity.
- lopis 1y agoWhile Perplexity is able to show sources for the information it shows, the language part, and the body of text upon which it was trained, is a black box, and sources are not given, nor typically desirable as a user.
- lopis 1y ago> Artists and creatives can cite their inspirational sources Even humans have a lot of internalized unconscious inspirational sources, but I get your point.
- msabalau 1y agoAI Boosters can suggest whatever nonsense strikes their fancy, and creatives can give into fear for no reason, but the best estimates we have from the BLS is that the there are careers and ongoing demand for artists, writers, photographers. Regardless, deep learning models are valuable because they generalize within the training data to uncover patterns and features and relationships that are implicit, rather (simply) present with the data. While they can return things that happen to be within the training set, there is no reason to believe that any particular output is literally found there or is something that could be attributable, or that a human would ever attribute. Human artists also make meaning from the broad texture of their life experiences and general diffuse unattributable experience of culture. Sure, this is something a random artist is unlikely to know, but if they are simply refusing to pick up a useful tools that can't give credit--say avoiding LLMs for brainstorming, or generative selection tools for visual editing, or whatever, their particular careers will be harmed by their incurious sentimentality, and other human artists will thrive because they know that tools are just tools, and it is the humans using the tools that make meaning that people care about.
- ijk 1y agoA difficult, but not intractable problem: OLMoTrace claims to be able to trace from output to training data in seconds [1]. Notably, it can do this because OLMo itself was intentionally designed to be open and transparent [2]; it was trained on 4.6 trillion tokens of entirely open data (which you can download yourself) [3]. There's nothing stopping Meta or OpenAI from creating a similar tool, other than the obvious detail of that showing their exact training data. [1] https://arxiv.org/abs/2504.07096 https://arxiv.org/abs/2504.07096 [2] https://allenai.org/blog/olmotrace https://allenai.org/blog/olmotrace [3] https://huggingface.co/datasets/allenai/olmo-mix-1124 https://huggingface.co/datasets/allenai/olmo-mix-1124
- stego-tech 1y agoI love it! Keeping this in my back pocket the next time someone claims that keeping accounting of training data and sourcing it isn't feasible or technically possible.
- immibis 1y agoWell, no, because all those file-sharing users used to get fined $250,000 or whatever, which is obviously much greater than the amount they would have paid for whatever they downloaded.
- internetter 1y ago> Meta did pirate the works but may be entitled to use them under fair use What fair use? Were the books promised to them by god or something?
- TimPC 1y agoFair use allows for certain uses of copyrighted works without a specific license for those works. One of the major criterion is how transformative the work is and an LLM model is very different from the original work so it seems likely that criterion at least is met.
- SideburnsOfDoom 1y ago> an LLM model is very different from the original work True, but not the only relevant thing. If the output of the LLM is "not very different from the original work" then the output could be the infringement. Putting a hypercomplex black box between the source work and the plagiarised output does not in itself make it "not infringing". The "LLM output as a service" business is then based on selling something based other people's work, that they do not have rights to. It's falling for misdirection, "pay no attention to the LLM behind the curtain" to think otherwise.
- Filligree 1y agoThe output of the LLM is very different from the original, though. It’s hard to look at this and claim it isn’t.
- SideburnsOfDoom 1y ago> The output of the LLM is very different from the original, though I will disagree with that characterisation. IMHO: In some cases no, it's not different, there are clear lines from inputs to output. In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. But I think that this is legally undecided and won't be decided by you or me, and it is going to be a more interesting and relevant question than "is the LLM model is very like the original work", which it clearly isn't. That's like asking "is this typewriter like this novel?" It can't be, but the words that came out of it could be.
- SideburnsOfDoom 1y agoFirstly, no kidding, of course it's "illegal" and "Piracy". Secondly, there's an argument that the infringement happens only when the LLM produces output based in part of whole on the source material. In other words, training a model is not infringing in itself. You could "research" with it. But selling the output as "from your model" is highly suspect. Your business is then based on selling something based other people's work, that you do not have rights to.
- gabriel666smith 1y agoI think there's a really fundamental misunderstanding of the playing field in this case. (Disclaimer that my day job is 'author', and I'm pro-piracy.) We need to frame this case - and ongoing artist-vs-AI-stuff -using a pseudoscience headline I saw recently: 'average person reads 60k words/day'. I won't bother sourcing this, because I don't think it's true, but it illustrates the key point: consumers spend X amount of time/day reading words. > It seems like the authors are setting up for failure by making the case about whether the AI generation hinders the market for books. AI book writing is such a tiny segment what these models do that if needed Meta would simply introduce guard rails to prevent copying the style of an author and continue to ingest the books. and from the article: > When he turned to the authors’ legal team, led by high-profile attorney David Boies, Chhabria repeatedly asked whether the plaintiffs could actually substantiate accusations that Meta’s AI tools were likely to hurt their commercial prospects. “It seems like you’re asking me to speculate that the market for Sarah Silverman’s memoir will be affected,” he told Boies. “It’s not obvious to me that is the case.” The market share an author (or any other artist type) is competing with for Meta is not 'what if an AI wrote celebrity memoirs?'. Meta isn't about to start a print publishing division. Authors are competing with Meta for 'whose words did you read today?' Were they exclusively Meta's - Instagram comments, Whatsapp group chat messages, Llama-generated slop, whatever - or did an author capture any of that share? The current framing is obviously ludicrous; it also does the developers of LLMs (the most interesting literary invention since....how long ago?) a huge disservice. Unfortunately the other way of framing it (the one I'm saying is correct) is (probably) impossible to measure (unless you work for Meta, maybe?) and, also, almost equally ridiculous.
- kazinator 1y agoDo you not understand that "fair use" is not some copyright free-for-all which lets you use works wholesale without attribution as if they were suddenly public domain? To make fair use of a book's passage, you have to cite it. The except has to be reasonably small. Without fair use, it would not be possible to write essays and book reviews that give quotes from books. That's what it's for. Not for having a machine read the whole book so it can regurgitate mashups of any part of it without attribution. Making a parody is a kind of fair use, but parodies are original expression based on a certain structure of the work.
- danaris 1y ago> To make fair use of a book's passage, you have to cite it. That's not true. That's what's required for something not to be plagiarism, not for something not to be copyright infringement. Fair use is not at all the same as academic integrity, and while academic use is one of the fair use exceptions, it's only one. The most you would have to do with any of the other fair use exceptions is credit where you got the material (not cite individual passages), because you're not necessarily even using those passages verbatim.
- kazinator 1y agoIf your "fair use" is - of a commercial nature; - plagiarism; - substantially large (e.g. whole work); you're not on good legal footing.
- danaris 1y agoI mean...sure? But that's not what you were saying. You said, as a blanket statement, that fair use requires citing the passage, which is not true. Fair use and plagiarism are related, but they are two separate things. Especially when talking about the legalities of things, as we are here, it's vital to be clear and accurate about what specific legal issues are under discussion. Facebook isn't claiming fair use to write academic papers about the books they're taking parts from. They're claiming fair use to feed them into LLM training. If that is a usage that is deemed to fall under fair use, then it won't require specific citation, even if it requires attribution in the more general sense (ie, crediting all the works you fed into your word-chipper), any more than you're required to cite specific passages when you're making a wholesale parody of a copyrighted work (also fair use) or writing a fanfic based on it (also fair use).
- apercu 1y ago> but may be entitled to use them under fair use. Why? Was it legal for me to download copyrighted songs from Limewire as "fair use"? Because a few people were made examples of. I'm a musician, so 80% of the music I listen to is for learning so it's fair use, right? ;)
- Filligree 1y ago> I'm a musician, so 80% of the music I listen to is for learning so it's fair use, right? ;) I would be happy with that outcome. I’m a fanfiction writer, and a lot of the stories I read are very much for learning. ;-)
- BrawnyBadger53 1y agoIf the result of this becomes that substantial remixes and fanfiction can be commercialized without permission from authors then I am happy. This stuff should have been fair use to begin with. Granted it probably already is fair use but because of the way copyright is enforced online it is effectively banned regardless.
- lukeschlather 1y agoI don't believe anyone was ever penalized for downloading only uploading which seems like a pretty similar principle to what the judge is saying here.
- deleted 1y ago[deleted]
- fngjdflmdflg 1y agoThis is why Meta didn't seed.[0] [0] https://torrentfreak.com/meta-says-it-made-sure-not-to-seed-any-pirated-books/ https://torrentfreak.com/meta-says-it-made-sure-not-to-seed-...
- sillysaurusx 1y agoHeh. People were penalized for merely creating search engines that happened to link to songs. Supposedly the RIAA accepted the offer of a 20-something’s life savings, but only if they switched their major from CS to something else. I believe it, having witnessed those times.
- aurizon 1y agoYes, current AI video/text product is inferior at this time. Youtube is full of all genres of inferior products - at this time! The ramp of improvement is quite steeply pointing upwards. This is retrospective of the days of spinning jennies and knitting/weaving machines that soon made manual products un-economic - that said, excellent craft/art product endured on a smaller scale. AI is also taking a toll on the movie arts, staring at the low end and climbing the same incremental improvement rungs. All the special effects(SFX) are in a similar boat. Prop rentals are hit hard. 100 high res photos of an old Studio Tv camera - all angles/sizes/lighting can be added to an AI prop library and with a green screen insert the prop can manifest as a true object in any aspect. There can be many. It still takes people to cull the hallucinations - a declining problem. Same with actors. They can be patterned after a famous actor - with likeness fees, or created de-novo. All the classic aspects of a studio production suffer the same incremental marginalisation - in 5 years = what will remain? - what new tech will emerge? I feel that many forks will emerge, all fighting for a place in the sun = some will be weeded out, some will flower - but at a very high pace. The old producers/directors/writers - the whole panoply of what makes a major studio will be scattered like dried bread crumbs,
- onlyrealcuzzo 1y ago> I also don’t think AI generated fiction is anywhere near high quality enough to substantially reduce the market for the original author. Legal cases are often based on BS, really an open form of extortion. The plaintiffs might've been hoping for a settlement. Meta could pay $xM+ to defend itself. Maybe they thought Meta would be happy to pay them $yM to go away. The reality is, there's very little Meta couldn't just find a freely available substitute for if it had to, it might just take a little more digging on their end. The idea that any one individual or small group is so valuable that can hold back LLMs by themselves is ridiculous. But you'll find no end to people vain enough to believe themselves that important.