11 ms·
Judge said Meta illegally used books to build its AI
- labrador 1y ago"Chhabria is cutting through the moral noise and zeroing in on economics. He doesn't seem all that interested in how Meta got the data or how “messed up” it feels—he’s asking a brutally simple question: Can you prove harm?" https://archive.is/Hg4Xr https://archive.is/Hg4Xr
- trinsic2 1y ago> Can you prove harm? Where was this argument when Napster was being sued?
- TimPC 1y agoI think the headline is a bit misleading. Mets did pirate the works but may be entitled to use them under fair use. It seems like the authors are setting up for failure by making the case about whether the AI generation hinders the market for books. AI book writing is such a tiny segment what these models do that if needed Meta would simply introduce guard rails to prevent copying the style of an author and continue to ingest the books. I also don’t think AI generated fiction is anywhere near high quality enough to substantially reduce the market for the original author.
- stego-tech 1y agoThe problem is that "harm" as defined by copyright law is strictly limited to loss of sales due to breach of that copyright; it makes no allowment (that I know of) to livelihoods lost by the theft of the work indefinitely, as AI boosters suggest their tools can do (replace people). The way this court case is going, it's an uphill battle for the plaintiffs to prove concrete harm in that very narrow context, when the real harm is the potential elimination of their future livelihoods through theft, rather than immediately tangible harms. As a (creative) friend of mine flatly said, they refuse to use an LLM until it can prove where it learned something from/cite its original source. Artists and creatives can cite their inspirational sources, while LLMs cannot (because their developers don't care about credit, only output) by design. To them, that's the line in the sand, and I think that's a reasonable one given that not a single creative in my circles has been cut payment from these multi-billion-dollar AI companies for the unauthorized use of their works in training these models.
- tedivm 1y agoGithub does give free copilot access to open source developers it considers important enough (which is a pretty low bar). While not the same as actually paying, it's the only example I can think of where the company that used people's copyrighted material actually gave something back to those people.
- _aavaa_ 1y agoWhat you’re describing is the Extend phase of Microsoft’s plan.
- mistrial9 1y agothe educated and erudite can wait in line near the Castle; every day due to the grace of our masters, unused bread from the master's table is available without prejudice. These people can have a fine life, and the people are fulfilled.
- bgwalter 1y agoThey want to track and utilize the new code that those developers are writing. And they want to keep them on GitHub. And they want to claim in potential lawsuits: "See, those developers themselves have used CoPilot, so they approve the copyright infringement."
- subscribed 1y agoYour friend might want to check Perplexity.
- lopis 1y agoWhile Perplexity is able to show sources for the information it shows, the language part, and the body of text upon which it was trained, is a black box, and sources are not given, nor typically desirable as a user.
- lopis 1y ago> Artists and creatives can cite their inspirational sources Even humans have a lot of internalized unconscious inspirational sources, but I get your point.
- internetter 1y ago> Meta did pirate the works but may be entitled to use them under fair use What fair use? Were the books promised to them by god or something?
- TimPC 1y agoFair use allows for certain uses of copyrighted works without a specific license for those works. One of the major criterion is how transformative the work is and an LLM model is very different from the original work so it seems likely that criterion at least is met.
- SideburnsOfDoom 1y ago> an LLM model is very different from the original work True, but not the only relevant thing. If the output of the LLM is "not very different from the original work" then the output could be the infringement. Putting a hypercomplex black box between the source work and the plagiarised output does not in itself make it "not infringing". The "LLM output as a service" business is then based on selling something based other people's work, that they do not have rights to. It's falling for misdirection, "pay no attention to the LLM behind the curtain" to think otherwise.
- Filligree 1y agoThe output of the LLM is very different from the original, though. It’s hard to look at this and claim it isn’t.
- SideburnsOfDoom 1y ago> The output of the LLM is very different from the original, though I will disagree with that characterisation. IMHO: In some cases no, it's not different, there are clear lines from inputs to output. In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. But I think that this is legally undecided and won't be decided by you or me, and it is going to be a more interesting and relevant question than "is the LLM model is very like the original work", which it clearly isn't. That's like asking "is this typewriter like this novel?" It can't be, but the words that came out of it could be.
- SideburnsOfDoom 1y agoFirstly, no kidding, of course it's "illegal" and "Piracy". Secondly, there's an argument that the infringement happens only when the LLM produces output based in part of whole on the source material. In other words, training a model is not infringing in itself. You could "research" with it. But selling the output as "from your model" is highly suspect. Your business is then based on selling something based other people's work, that you do not have rights to.
- gabriel666smith 1y agoI think there's a really fundamental misunderstanding of the playing field in this case. (Disclaimer that my day job is 'author', and I'm pro-piracy.) We need to frame this case - and ongoing artist-vs-AI-stuff -using a pseudoscience headline I saw recently: 'average person reads 60k words/day'. I won't bother sourcing this, because I don't think it's true, but it illustrates the key point: consumers spend X amount of time/day reading words. > It seems like the authors are setting up for failure by making the case about whether the AI generation hinders the market for books. AI book writing is such a tiny segment what these models do that if needed Meta would simply introduce guard rails to prevent copying the style of an author and continue to ingest the books. and from the article: > When he turned to the authors’ legal team, led by high-profile attorney David Boies, Chhabria repeatedly asked whether the plaintiffs could actually substantiate accusations that Meta’s AI tools were likely to hurt their commercial prospects. “It seems like you’re asking me to speculate that the market for Sarah Silverman’s memoir will be affected,” he told Boies. “It’s not obvious to me that is the case.” The market share an author (or any other artist type) is competing with for Meta is not 'what if an AI wrote celebrity memoirs?'. Meta isn't about to start a print publishing division. Authors are competing with Meta for 'whose words did you read today?' Were they exclusively Meta's - Instagram comments, Whatsapp group chat messages, Llama-generated slop, whatever - or did an author capture any of that share? The current framing is obviously ludicrous; it also does the developers of LLMs (the most interesting literary invention since....how long ago?) a huge disservice. Unfortunately the other way of framing it (the one I'm saying is correct) is (probably) impossible to measure (unless you work for Meta, maybe?) and, also, almost equally ridiculous.
- kazinator 1y agoDo you not understand that "fair use" is not some copyright free-for-all which lets you use works wholesale without attribution as if they were suddenly public domain? To make fair use of a book's passage, you have to cite it. The except has to be reasonably small. Without fair use, it would not be possible to write essays and book reviews that give quotes from books. That's what it's for. Not for having a machine read the whole book so it can regurgitate mashups of any part of it without attribution. Making a parody is a kind of fair use, but parodies are original expression based on a certain structure of the work.
- danaris 1y ago> To make fair use of a book's passage, you have to cite it. That's not true. That's what's required for something not to be plagiarism, not for something not to be copyright infringement. Fair use is not at all the same as academic integrity, and while academic use is one of the fair use exceptions, it's only one. The most you would have to do with any of the other fair use exceptions is credit where you got the material (not cite individual passages), because you're not necessarily even using those passages verbatim.
- kazinator 1y agoIf your "fair use" is - of a commercial nature; - plagiarism; - substantially large (e.g. whole work); you're not on good legal footing.
- danaris 1y agoI mean...sure? But that's not what you were saying. You said, as a blanket statement, that fair use requires citing the passage, which is not true. Fair use and plagiarism are related, but they are two separate things. Especially when talking about the legalities of things, as we are here, it's vital to be clear and accurate about what specific legal issues are under discussion. Facebook isn't claiming fair use to write academic papers about the books they're taking parts from. They're claiming fair use to feed them into LLM training. If that is a usage that is deemed to fall under fair use, then it won't require specific citation, even if it requires attribution in the more general sense (ie, crediting all the works you fed into your word-chipper), any more than you're required to cite specific passages when you're making a wholesale parody of a copyrighted work (also fair use) or writing a fanfic based on it (also fair use).
- apercu 1y ago> but may be entitled to use them under fair use. Why? Was it legal for me to download copyrighted songs from Limewire as "fair use"? Because a few people were made examples of. I'm a musician, so 80% of the music I listen to is for learning so it's fair use, right? ;)
- Filligree 1y ago> I'm a musician, so 80% of the music I listen to is for learning so it's fair use, right? ;) I would be happy with that outcome. I’m a fanfiction writer, and a lot of the stories I read are very much for learning. ;-)
- BrawnyBadger53 1y agoIf the result of this becomes that substantial remixes and fanfiction can be commercialized without permission from authors then I am happy. This stuff should have been fair use to begin with. Granted it probably already is fair use but because of the way copyright is enforced online it is effectively banned regardless.
- lukeschlather 1y agoI don't believe anyone was ever penalized for downloading only uploading which seems like a pretty similar principle to what the judge is saying here.
- deleted 1y ago[deleted]
- fngjdflmdflg 1y agoThis is why Meta didn't seed.[0] [0] https://torrentfreak.com/meta-says-it-made-sure-not-to-seed-any-pirated-books/ https://torrentfreak.com/meta-says-it-made-sure-not-to-seed-...
- sillysaurusx 1y agoHeh. People were penalized for merely creating search engines that happened to link to songs. Supposedly the RIAA accepted the offer of a 20-something’s life savings, but only if they switched their major from CS to something else. I believe it, having witnessed those times.
- aurizon 1y agoYes, current AI video/text product is inferior at this time. Youtube is full of all genres of inferior products - at this time! The ramp of improvement is quite steeply pointing upwards. This is retrospective of the days of spinning jennies and knitting/weaving machines that soon made manual products un-economic - that said, excellent craft/art product endured on a smaller scale. AI is also taking a toll on the movie arts, staring at the low end and climbing the same incremental improvement rungs. All the special effects(SFX) are in a similar boat. Prop rentals are hit hard. 100 high res photos of an old Studio Tv camera - all angles/sizes/lighting can be added to an AI prop library and with a green screen insert the prop can manifest as a true object in any aspect. There can be many. It still takes people to cull the hallucinations - a declining problem. Same with actors. They can be patterned after a famous actor - with likeness fees, or created de-novo. All the classic aspects of a studio production suffer the same incremental marginalisation - in 5 years = what will remain? - what new tech will emerge? I feel that many forks will emerge, all fighting for a place in the sun = some will be weeded out, some will flower - but at a very high pace. The old producers/directors/writers - the whole panoply of what makes a major studio will be scattered like dried bread crumbs,
- onlyrealcuzzo 1y ago> I also don’t think AI generated fiction is anywhere near high quality enough to substantially reduce the market for the original author. Legal cases are often based on BS, really an open form of extortion. The plaintiffs might've been hoping for a settlement. Meta could pay $xM+ to defend itself. Maybe they thought Meta would be happy to pay them $yM to go away. The reality is, there's very little Meta couldn't just find a freely available substitute for if it had to, it might just take a little more digging on their end. The idea that any one individual or small group is so valuable that can hold back LLMs by themselves is ridiculous. But you'll find no end to people vain enough to believe themselves that important.
- MikeIndyson 1y ago[flagged]
- ndiddy 1y agoThe title for this submission is somewhat misleading. The judge didn't make any sort of ruling, this is just reporting on a pretrial hearing. He also doesn't seem convinced as to how relevant downloading books from LibGen is to the case: > At times, it sounded like the case was the authors’ to lose, with [Judge] Chhabria noting that Meta was “destined to fail” if the plaintiffs could prove that Meta’s tools created similar works that cratered how much money they could make from their work. But Chhabria also stressed that he was unconvinced the authors would be able to show the necessary evidence. When he turned to the authors’ legal team, led by high-profile attorney David Boies, Chhabria repeatedly asked whether the plaintiffs could actually substantiate accusations that Meta’s AI tools were likely to hurt their commercial prospects. “It seems like you’re asking me to speculate that the market for Sarah Silverman’s memoir will be affected,” he told Boies. “It’s not obvious to me that is the case.” > When defendants invoke the fair use doctrine, the burden of proof shifts to them to demonstrate that their use of copyrighted works is legal. Boies stressed this point during the hearing, but Chhabria remained skeptical that the authors’ legal team would be able to successfully argue that Meta could plausibly crater their sales. He also appeared lukewarm about whether Meta’s decision to download books from places like LibGen was as central to the fair use issue as the plaintiffs argued it was. “It seems kind of messed up,” he said. “The question, as the courts tell us over and over again, is not whether something is messed up but whether it’s copyright infringement.”
- bgwalter 1y agoThe RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. Now that big capital wants to steal from individuals, big capital wins again. (Unrelatedly, has Boies ever won a high profile lawsuit? I remember him from the Bush/Gore recount issue, where he represented the Democrats.)
- kranke155 1y agoCopyright was invented (in its modern form) by corporations. It will be uninvented if need be for corporations.
- adingus 1y agoI'm wondering if authors are making the same mistakes that the music industry did with Napster and kazaa. Using AI has led to more book purchases for me. If I discover and enjoy a book via AI I'm more inclined to buy it. The cats out of the bag, so pet him.
- tacheiordache 1y agoJust look at the state of the music industry.
- mtlynch 1y agoCan you share more about how you sample books with AI?
- nickpsecurity 1y agoIt can tell you about authors, books, useful techniques, etc. If it cites references, that can generate page views on their site ir sales. It can also replace that, though, with AI supplier benefiting commercially.
- adingus 1y agoI don't really sample them, but if I want to know more about a subject I will normally ask for a book recommendation to go along with it.
- aprilthird2021 1y agoMost people are not buying more books because of AI and no court would entertain this logic truly.
- Mbwagava 1y agoWhether or not Meta wins this case, I'm never going to support any government that supports both LLMs and IP. Like we have to put up with IP despite having no clear value to a digital society but as soon as it becomes inconvenient it goes out the window? Nah, let's just trash the state and start over. It's going to take centuries to undo the damage wracked by IP-supported private enterprise. And now we also have to put up with fucking chatbots. This is the worst timeline.
- jMyles 1y agoThe good news is that the internet is, fundamentally and in a way that no legacy state can alter, not a place where IP is cognizable. You are free to copy bytes as you see fit, and the internet treats them identically whether they are random noise or whether a codec can turn them into music, film, books, or whatever inspires you. The problem is that some humans, justifying their behavior by claiming it as "official", may act out with violence against you if they (rightly or wrongly, that's important to note) perceive that your actions are causing the internet to copy bytes to which they object. Enduring nonviolence is likely yet ahead as consensus grows over the end of the legitimacy of these legacy states.
- labrador 1y agoI hope you don't think I'm snarky because I'm serious. If you're an American citizen you can homestead in Alaska and cut yourself off from all this if you like. edit: i'm serious. many americans would be much happier taking this option if they knew it existed. i may take it myself
- sillysaurusx 1y agoHomesteading is tremendously expensive, unfortunately. Most people can’t.
- labrador 1y agoI didn't know that, but in that case there are a lot of young men and women on HN who are financically successful, but are tremendously unhappy. That's the case for me when I looked into it 25 years ago.
- RajT88 1y ago> “What about the next Taylor Swift?” he asked, arguing that a “relatively unknown artist” whose work was ingested by Meta would likely have their career hampered if the model produced “a billion pop songs” in their style. I have this debate with a friend of mine. He's terrified of AI making all of our jobs obsolete. He's a brilliant musician and programmer both, so he's both enthused and scared. So let's go with the Swift example they use. Performance Artists have always tried to cultivate an image, an ideal, a mythos around the character(s). I've observed that as the music biz has gotten more tough, that the practice of selling merch at shows has ramped up. Social media is monetized now. There's been a big diversification in the effort to make more money from everything surrounding the music itself. So too will it be with artists. You're starting to see this already. Artists which got big not necessarily because of the music, but because of the weird cult of personality they built. One who comes to minds is Poppy, who ironically enough built a cult of personality around her being a fake AI bot... https://en.wikipedia.org/wiki/Poppy_(singer) https://en.wikipedia.org/wiki/Poppy_(singer) You've definitely got counter-examples like Hatsune Miku - but the novelty of Miku was because of the artificiality (within a culture that, like, really loves robots and shit). AI pop stars will undoubtedly produce listenable records and make some money, but I don't expect that they will be able to replace the experience of fans looking for a connection with an artist. Watch the opening of a Taylor Swift concert, and you'll probably get it.
- atrus 1y agoI think that argument is further hampered (taylor being an exception) by the fact that most pop stars already don't write their own songs. If people like Max Martin can pump out multiple hit songs for multiple groups, it kinda shows that who wrote the song doesn't matter. Has making music for a living ever not been tough?
- RajT88 1y ago> Has making music for a living ever not been tough? Fair. > I think that argument is further hampered (taylor being an exception) by the fact that most pop stars already don't write their own songs. That accounts for the big artists on the radio (yes some people listen to that). But, what about everyone else? I would posit that most artists (the one-hit wonders, the ones without radio success, etc.) write their own songs. It seems like there's such acts who make a go of it just fine, who write their own songs and really nail the connection with fans. I would point to a regional band near me: Mr. Blotto.
- steele 1y agoIn a just world, this would shutter the organization.
- openplatypus 1y agoWhat is this just world you are talking about?
- kazinator 1y agoThe legal system is not going to be kind to the AI hucksters. Why? Because, quite stupidly and counterproductively, they have stepped on its toes by claiming that AI can replace lawyers. On top of that, there have been incidents of lawyers getting in hot water for generating slop instead of doing their work. So, this isn't just some distant, abstract tech issue for the lawyers and judges, like whether APIs should be copyrighted. If you're in any kind of business, you generally want these people to be on your side. Oopsies!
- throwacct 1y agoOnly Facebook?!!
- Workaccount2 1y agoLet me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguously illegal. Meta is charged with doing the latter, but it seems the plaintiffs want to also tie in the former.
- Lerc 1y agoI'm not sure if Meta did anything illegal in 2. either. I thought the copyright infringement was by the people who provided the copyrighted material when they did not have the rights to do so. I may be wrong on this, but it would seem a reasonable protection for consumers in general. Meta is hardly an average consumer, but I doubt that matters in the case of the law. Having grounds to suspect that the provider did not have the rights might though.
- CyberMacGyver 1y agoThe source being illegal doesn’t make your use legal. Infact one could argue that it’s equally illegal or worse since a corporation knowingly engaged in illegal activity.
- Lerc 1y agoWell obviously, but the converse is also true. The source being illegal doesn't make your use illegal. In the eyes of the law it doesn't matter if something is better or worse, it's the law, it shouldn't be confused with morality. An act is illegal if the law says it is and isn't if it isn't. Just by being involved as a party does not make you culpable. Murderers are criminals, the murdered, less so. Choosing to be a party might not make you culpable. You may be an active participant but unaware of the law breaking (being defrauded). Or the law may explicitly state that you can engage with people committing criminal acts and reap the benefits so long as you don't break those laws (or encourage them to be broken) yourself. Some forms of journalism are protected in this way. Ultimately to have a case you have to state. 1. What law was broken 2. How an action by a party is in violation of that law. 3. That the action actually happened. The largest problem with this case is not that 3. is in doubt but showing which 1. and 2. they are talking about.
- ebfe1 1y agoAnd this is how Chinese model will win in long term, perhaps... They will be trained on everything and anything without consequences and we will all use it because these models are smarter (except for area like Chinese history and geography). I don't have the right answer on what can be done here to protect copyright or rather contributing back to authors of a paper without all these millions dollar wasted in lawsuits.
- aprilthird2021 1y agoThere's no winning though. There's no real moat when it comes to AI remember. There will be tons of models of similar, squishy types of unique attributes (squishy meaning it works great sometimes and not other times, and that's just normal). And it will mostly be decided which to use based on cost and compliance.
- ryandrake 1y agoAI hucksters vs. the Copyright Cartel. When two evil villains fight, who do you root for? Here's hoping they somehow destroy each other.
- thomastjeffery 1y agoI can only root for them both to lose. Letting Meta launder copyrighted works to make billions, while threatening the rest of us over the most trivial derivative work, sounds like the worst outcome to me. Copyright is a mistake. It demands that we compete instead of collaborate. LLMs don't provide enough utility to deserve special treatment in these circumstances. If anyone can infringe copyright, then everyone should be able to.
- probably_wrong 1y agoYou can always root for the lawyers.
- pessimizer 1y agoLast time they fought, it was google vs. the publishers, and it resulted in the scanning and archiving of all of those books in the first place. Neither of them died, though, both parties just kept all the books from the public and used them for their own purposes, while normal people had to squirrel them away and trade them illegally. It's the Tech Cartels vs. the Copyright Trolls. It'll end up as a romance.
- akomtu 1y agoThe copyright cartel is going to lose because it represents the dying old world order of many competing enclaves. AI isn't just a sloppy text generator, it's the new ideology of forced uniformity that permits no boundaries. So no copyright cartels.
- caminanteblanco 1y agoI feel like this submission does a disservice by changing the title from the article's. It is misleading, and implied that the judge has already given a ruling, when they have not.
- dragonwriter 1y agoThis is the source headline, but it is pure clickbait; the judge absolutely did not say that in any of the quotes in the article; in the hearing on both parties motions for partial sunmary judgement, he both said that would be the case if the plaintiffs proved certain facts and raised doubts that they have the evidence to prove them.
- option 1y agothis is a huge issue AI companies in China do not have. the law must adjust now.
- deleted 1y ago[deleted]
- zoobab 1y agoJust went to the public library and read a book to train my brain without permissions from the authors.
- lern_too_spel 1y agoDid you then distribute copies of your brain to other people who used them to reproduce the copyrighted works verbatim? https://www.patronus.ai/blog/introducing-copyright-catcher https://www.patronus.ai/blog/introducing-copyright-catcher
- jayd16 1y agoSo to the legal peanut gallery here... What is the substantive difference between training a model locally using these works that are presumably pulled in from some database somewhere and Napster, for example? Would a p2p network for sharing of copyrighted works be legal if the result is to train a model? What if I promise the model can't reproduce the works verbatim?
- penguin_booze 1y agoGood. Now do OpenAI.
- TrnsltLife 1y agoReading the books changes the weights of the neural network. If ruled illegal, wouldn't it also become illegal for a human to read an illegally downloaded book? So far, I thought just redistribution was illegal. Will the neural network (LLM) itself become illegal? Will its outputs be deemed illegal? If so, do humans who have read an illegally downloaded book become illegal? Do their creative outputs become illegal?
- delecti 1y agoBooks are sold for the purpose of people reading them, including all the normal consequences that happen from a person reading a book. AI training being analagous to that doesn't unlock some cheat code that makes it legal, or reading books illegal. And it might indeed be found legal, but not for that reason.
- codedokode 1y agoIt's typical double standards policy: Google and Github remove links to pirated material (and pirated material itself) so that ordinary folks cannot download it for free, but when Zuckerberg downloads gigabytes of pirated material without paying, it's ok. The legal system doesn't want to put an ordinary folk and Zuckerberg at the same level. Also I read that ordinary folks have been arrested for filming in the cinema even if they did not redistribute the video (due to being arrested). Again, it is unfair why they get arrested and Zuckerberg doesn't.
- granzymes 1y agoTitle seems misleading after reading the article.
- gtowey 1y agoIt's mind blowing to me that the court might deny the right of the authors to control licensing of this kind of usage of their work.
- deleted 1y ago[deleted]
- jwatte 1y agoIf I put something up for anyone to read on the internet. And someone reads it on the internet. I can't really control that, right? Now, if someone makes an infringing use of the thing I put up on the internet. Then I have some kind of recourse, at least through the courts, if I have a lot of money to pay lawyers. But if someone makes a fair use of the thing I put up on the internet, then I don't have any recourse, because that's the way the law works. As far as I understand it, using data as input data to a machine learning model that substantially transforms and does not duplicate the input data is currently believed to be fair use. So, the training use of freely available data seems pretty straightforward that authors can't control when they make it freely available. It seems like Facebook made use of data that wasn't freely available, though -- ebook rip library type stuff. That's the bit I think they could be in trouble for. But that's just a plain-old "Napster" style copyright question, as far as I understand it. The lawyer's argument that Llama "obliterates the market" for written works seems weak. I, and anyone I know, put down AI slop fiction before the first paragraph is done, because it's not the same thing as real fiction.
- moregrift 1y agoThis is the main reason why Chinese AI will be better than Western AI in the long term - Chinese companies can train on higher quality dataset (all the copyrighted books in the world)
- deleted 1y ago[deleted]
- codr7 1y agoMeta did what they always do, whatever they think they can get away with.
- nottorp 1y agoOnly Meta?
- codedokode 1y agoIf it is a "fair use" to run the business using pirated books, does it mean that it is "fair use" to use pirated software as long as you don't distribute it? Why pay for copyrighted works if Zuckerberg downloads them for free?
- terbo 1y agoMeanwhile Chinese models are uncensored, trained on everything they can get, and outperform restricted models ..
- ineedasername 1y agoI simply don't think that the copyright IP framework as it exists can be applied to training on this scale. Or, if it can, the relative value of any specific author/content creator's work is deminimis. When the scale is a significant portion of all human text output ever, I don't think we're in the realm of any prior model. This is now something closer to how society attempts to approach natural resources like land, frequency bands, utility right-of-way, etc. I think this is the direction that laws and legislation should look to go. Or maybe not, I don't claim to have the answer, only that existing models are inadequate.
- PeterStuer 1y agoI get both sides of this debate. However, claiming llama is not a 'substantial transformation' of the information used to build it seems untennable. The complaint feels to me more like the paint factory claiming rights to the paintings you created with it's paint, rather than a classic pirate DVD copier that just resells copies. Maybe a midway could be some Google Books like solution where you can still find anything but where the output is restricted to just substantial fragments and not complete verbatim chapters? I do not believe people use llama to 'read published books on the cheap'.