10 ms·
Well, firing someone for this is super weird. It seems like an attempt to censor an interpretation of the law that: 1. Criticizes a highly useful technology 2.
by mattxxx 1y ago
Well, firing someone for this is super weird. It seems like an attempt to censor an interpretation of the law that:
1. Criticizes a highly useful technology
2. Matches a potentially-outdated, strict interpretation of copyright law
My opinion: I think using copyrighted data to train models for sure seems classically illegal. Despite that, Humans can read a book, get inspiration, and write a new book and not be litigated against. When I look at the litany of derivative fantasy novels, it's obvious they're not all fully independent works.
Since AI is and will continue to be so useful and transformative, I think we just need to acknowledge that our laws did not accomodate this use-case, then we should change them.
- madeofpalk 1y ago> Humans can read a book, get inspiration, and write a new book and not be litigated against Humans get litigated against this all the time. There is such thing as, charitably, being too inspired. https://en.wikipedia.org/wiki/List_of_songs_subject_to_plagiarism_disputes https://en.wikipedia.org/wiki/List_of_songs_subject_to_plagi...
- jrajav 1y agoIf you follow these cases more closely over time you'll find that they're less an example of humans stealing work from others and more an example of typical human greed and pride. Old, well established musicians arguing that younger musicians stole from them for using a chord progression used in dozens of songs before their own original, or a melody on the pentatonic scale that sounds like many melodies on the pentatonic scale do. It gets ridiculous. Plus, all art is derivative in some sense, it's almost always just a matter of degree.
- FireBeyond 1y agoTo the point that Billy Joel "famously" credited the songwriter for one of his songs ("This Night") as "Billy Joel, Ludwig van Beethoven".
- johnnyanmac 1y ago> art is derivative in some sense, it's almost always just a matter of degree. Yes, that's why we judge on a case by case basis. The line is blurry. I think when you're storing copies of such assets in your database that you're well past the line, though.
- ActionHank 1y agoAssuming this means copyright is dead, companies will be vary upset and patents will likely follow. The hold US companies have on the world will be dead too. I also suspect that media piracy will be labelled as the only reason we need copyright, an existing agency will be bolstered to address this concern and then twisted into a censorship bureau.
- palmotea 1y ago[flagged]
- jobigoud 1y agoWe are talking about the rights of the humans training the models and the humans using the models to create new things. Copyright only comes into play on publication. It's only concerned about publication of the models and publication of works. The machine itself doesn't have agency to publish anything at this point.
- MyOutfitIsVague 1y agoIt's not only publication, otherwise people wouldn't be able to be successfully sued for downloading and consuming copyrighted content, it would only be the uploaders who get into trouble.
- HappMacDonald 1y agoDo you have any links to cases where people were sued for downloading and consuming content without also uploading (eg, bittorent), hosting, sharing the copyrighted works, etc?
- MyOutfitIsVague 1y agoThere were the famous napster cases, the kids and old ladies that got sued by the RIAA for using limewire to download some music. There is also the fact that copyright holders will pressure your ISP into sending threatening letters and shutting off your Internet for piracy, even without you seeding. I haven't gotten the impression that you are in the clear for pirating as long as you don't distribute.
- lavezzi 1y agoThere's tonnes, this is a baffling question.
- bgwalter 1y ago
- timdiggerm 1y agoOr we could acknowledge that something could be a bad idea, despite its utility
- stevenAthompson 1y agoDoing a cover song requires permission, and doing it without that permission can be illegal. Being inspired by a song to write your own is very legal. AI is fine as long as the work it generates is substantially new and transformative. If it breaks and starts spitting out other peoples work verbatim (or nearly verbatim) there is a problem. Yes, I'm aware that machines aren't people and can't be "inspired", but if the functional results are the same the law should be the same. Vaguely defined ideas like your soul or "inspiration" aren't real. The output is real, measurable, and quantifiable and that's how it should be judged.
- toast0 1y ago> Doing a cover song requires permission, and doing it without that permission can be illegal. I believe cover song licensing is available mechanically; you don't need permission, you just need to follow the procedures including sending the licensing fees to a rights clearing house. Music has a lot of mechanical licenses and clearing houses, as opposed to other categories of works.
- stevenAthompson 1y ago> you don't need permission, you just need to follow the procedures Those procedures are how you ask for permission. As you say, it usually involves a fee but doesn't have to.
- toast0 1y ago(in the US) Mechanical licenses are compulsory; you don't need permission, you can just follow the forms and pay the fees set by the Copyright Royalty Board (appointed by the Librarian of Congress). You can ask the rightsholder to negotiate a lower fee, but there's no need for consent of the rightsholder if you notify as required (within 30 days of recording and before distribution) and pay the set fees.
- stevenAthompson 1y ago
- jeroenhd 1y agoPirating movies is also useful, because I can watch movies without paying on devices that apps and accounts don't work on. That doesn't make piracy legal, even though I get a lot of use out of it. Also, a person isn't a computer so the "but I can read a book and get inspired" argument is complete nonsense.
- Workaccount2 1y agoIt's only complete non-sense if you understand how humans learn. Which we don't. What we do know though is that LLMs, similar to humans, do not directly copy information into their "storage". LLMs, like humans, are pretty lossy with their recall. Compare this to something like a search indexed database, where the recall of information given to it is perfect.
- zelphirkalt 1y agoWell, you don't get to pick and choose in which situations an LLM is considered similar to a human being and in which not. If you argue that it similarly to a human is lossy, well let's go ahead and get most output checked by organizations and courts for violations of the law and licenses, just like human work is. Oh wait, I forgot, LLMs are run by companies with too much cash to successfully sue them. I guess we just have to live with it then, what a pity.
- philipkglass 1y agoThere are a couple of ways to theoretically prevent copyright violations in output. For closed models that aren't distributed as weights, companies could index perceptual hashes of all the training data at a granular level (like individual paragraphs of text) and check/retry output so that no duplicates or near-duplicates of copyrighted training data ever get served as a response to end users. Another way would be to train an internal model directly on published works, use that model to generate a corpus of sanitary rewritten/reformatted data about the works still under copyright, then use the sanitized corpus to train a final model. For example, the sanitized corpus might describe the Harry Potter books in minute detail but not contain a single sentence taken from the originals. Models trained that way wouldn't be able to reproduce excerpts from Harry Potter books even if the models were distributed as open weights.
- vessenes 1y agoThank you - a voice of sanity on this important topic. I understand people who create IP of any sort being upset that software might be able to recreate their IP or stuff adjacent to it without permission. It could be upsetting. But I don't understand how people jump to "Copyright Violation" for the fact of reading. Or even downloading in bulk. The Copyright controls, and has always controlled, creation and distribution of a work. In the nature even of the notice is embedded the concept that the work will be read. Reading and summarizing have only ever been controlled in western countries via State's secrets type acts, or alternately, non-disclosure agreements between parties. It's just way, way past reality to claim that we have existing laws to cover AI training ingesting information. Not only do we not, such rules would seem insane if you substitute the word human for "AI" in most of these conversations. "People should not be allowed to read the book I distributed online if I don't want them to." "People should not be allowed to write Harry Potter fanfic in my writing style." "People should not be allowed to get formal art training that involves going to museums and painting copies of famous paintings." We just will not get to a sensible societal place if the dialogue around these issues has such a low bar for understanding the mechanics, the societal tradeoffs we've made so far, and is able to discuss where we might want to go, and what would be best.
- jasonlotito 1y ago> But I don't understand how people jump to "Copyright Violation" for the fact of reading. The article specificaly talks about the creation and distribution of a work. Creation and distribution of a work alone is not a copyright violation. However, if you take in input from something you don't own, and genAI outputs something, it could be considered a copyright violation. Let's make this clear; genAI is not a copyright issue by itself. However, gen AI becomes an issue when you are using as your source stuff you don't have the copyright or license to. So context here is important. If you see people jumping to copyright violation, it's not out of reading alone. > "People should not be allowed to read the book I distributed online if I don't want them to." This is already done. It's been done for decades. See any case where content is locked behind an account. Only select people can view the content. The license to use the site limits who or what can use things. So it's odd you would use "insane" to describe this. > "People should not be allowed to write Harry Potter fanfic in my writing style." Yeah, fan fiction is generally not legal. However, there are some cases where fair use covers it. Most cases of fan fiction are allowed because the author allows it. But no, generally, fan fiction is illegal. This is well known in the fan fiction community. Obviously, if you don't distribute it, that's fine. But we aren't talking about non-distribution cases here. > "People should not be allowed to get formal art training that involves going to museums and painting copies of famous paintings." Same with fan fiction. If you replicate a copyrighted piece of art, if you distribute it, that's illegal. If you simply do it for practice, that's fine. But no, if you go around replicating a painting and distribute it, that's illegal. Of course, technically speaking, none of this is what gen AI models are doing. > We just will not get to a sensible societal place if the dialogue around these issues has such a low bar for understanding the mechanics I agree. Personifying gen AI is useless. We should stick to the technical aspects of what it's doing, rather than trying to pretend it's doing human things when it's 100% not doing that in any capacity. I mean, that's fine for the the layman, but anyone with any ounce of technical skill knows that's not true.
- ceejayoz 1y ago> Despite that, Humans can read a book, get inspiration, and write a new book and not be litigated against. You're still not gonna be allowed to commercially publish "Hairy Plotter and the Philosophizer's Rock".
- WesolyKubeczek 1y agoNo, but you are most likely allowed to commercially publish "Hairy Potter and the Philosophizer's Rock", a story about a prehistoric community. The hero is literally a hairy potter who steals a rock from a lazy deadbeat dude who is pestering the rest of the group with his weird ideas.
- zelphirkalt 1y agoNot sure what you are getting at?
- anigbrowl 1y agoYou are if it's parody, cf 'Bored of the Rings'.
- otabdeveloper4 1y ago[flagged]
- regularjack 1y agoThen they need to be changed for everyone and not just AI companies, but we all know that ain't happening.
- zelphirkalt 1y agoThe law covers these cases pretty well, it is just that the law has very powerful extremely rich adversaries, whose greed has gotten the better of them again and again. They could use work released sufficiently long ago to be legally available, or they could take work released as creative commons, or they could run a lookup, to make sure to never output verbatim copies of input or outputs, that are within a certain string editing distance, depending on output length, or they could have paid people to reach out to all the people, whose work they are infringing upon. But they didn't do any of that, of course, because they think they are above the law.
- nadermx 1y agoI'm confused, so you're saying its illegal? Because last I checked it's still in the process of going through the courts. And need we forget that copyright's purpose is to advance the arts and sciences. Fair use is codified into law, which states each case is seen on a use by use basis, hence the litigation to determine if it is in fact, legal.
- mdhb 1y agoIt’s so fucking obviously illegal when you think about it rationally for more than a few seconds. We aren’t even talking about “fair use” we are talking about how it works in practice which was Meta torrenting pirated books, never paying anyone a cent and straight up stealing the content at scale.
- Intralexical 1y agoA test to apply here: If you or I did this, would it be illegal? Would we even be having this conversation? The law is supposed to be impartial. So if the answer is different, then it's not really a law problem we're talking about.
- nadermx 1y agoThe fact you are even using the word stealing, is telling to your lack of knowledge in this field. Copyright infringement is not stealing[0]. The propaganda of the copyright cartel has gotten to you. [0] https://en.wikipedia.org/wiki/Dowling_v._United_States_(1985) https://en.wikipedia.org/wiki/Dowling_v._United_States_(1985...
- franczesko 1y ago> Piracy refers to the illegal act of copying, distributing, or using copyrighted material without authorization. It can occur in various forms Professing of IP without a license AND offering it as a model for money doesn't seem like an unknown use-case to me
- deleted 1y ago[deleted]
- apercu 1y ago>Despite that, Humans can read a book, get inspiration, and write a new book and not be litigated against. Corporations are not humans. (It's ridiculous that they have some legal protections in the US like humans, but that's a different issue). AI is also not human. AI is also not a chipmunk. Why the comparison?
- SilasX 1y ago>My opinion: I think using copyrighted data to train models for sure seems classically illegal. Despite that, Humans can read a book, get inspiration, and write a new book and not be litigated against. When I look at the litany of derivative fantasy novels, it's obvious they're not all fully independent works. Huh? If you agree that "learning from copyrighted works to make new ones" has traditionally not been considered infringement, then can you elaborate on why you think it fundamentally changes when you do it with bots? That would, if anything, seem to be a reversal of classic copyright jurisprudence. Up until 2022, pretty much everyone agreed that "learning from copyrighted works to make new ones" is exactly how it's supposed to work, and would be horrified at the idea of having to separately license that. Sure, some fundamental dynamic might change when you do it with bots, but you need to make that case in an enforceable, operationalized way.
- bitfilped 1y agoSorry but AI isn't that useful and I don't see it becoming any more useful in the near term. It's taken since ~1950 to get LLMs working well enough to become popular and they still don't work well.
- dns_snek 1y agoThe problem with this kind of analysis is that it doesn't even try to address the reasons why copyright exists in the first place. This belief that training LLMs on content without permission should be allowed is incompatible with the belief that copyright is useful, you really have to pick a lane here. Go back to the roots of copyright and the answers should be obvious. According to the US constitution, copyright exists "To promote the Progress of Science and useful Arts" and according to the EU, "Copyright ensures that authors, composers, artists, film makers and other creators receive recognition, payment and protection for their works. It rewards creativity and stimulates investment in the creative sector." If I publish a book and tech companies are allowed to copy it, use it for "training", and later regurgitate the knowledge contained within to their customers then those people have no reason to buy my book. It is a market substitute even though it might not be considered such under our current copyright law. If that is allowed to happen then investment will stop and these books simply won't get written anymore.
- hochstenbach 1y agoHumans are not allowed to do what AI firms want to do. That was one of the copyright office arguments: a student can't just walk into a library and say "I want a copy of all your books, because I need them for learning". Humans are also very useful and transformative.
- p0w3n3d 1y agoit's funny how a law becomes potentially-outdated only when big corporations want to violate in on a global scale. As a private person I no longer feel incentivised to create new content online because I think that all I create will eventually be stolen from me...