33 ms·
Things are about to get worse for generative AI
- Baldbvrhunter 3y agoI imagine the argument might be like this: I hire a session musician to play on my new single, paying him $100. I record the whole session. I ask him to play the opening to "Stairway to Heaven" and he does so. "Well, I can't use that as a sample without paying" "Ok play something like Jimmy Page" "Hmm, still sounds like Stairway to Heaven" "Ok, try and sound less like Stairway to Heaven but in that style" "Great, I'll use that one" and I release my song and get $5,000 in royalties. Should I be sued for infringement, or the guitarist? The problem, I suppose, is that if I had said "play something like 70s prog rock" and he played "Stairway to Heaven" and I didn't know what it was and said "great, I'll use that". Should I be sued for infringement, or the guitarist?
- earthnail 3y agoYou are sued for infringement if you are the rightsholder. You need an agreement with the guitarist about rights. The default agreement for session musicians is that you pay them in return for their rights. It’s like a software engineering contractor. The contractor gets paid, the IP of their work is owned by the company.
- foobazgt 3y agoBut it's not like that. Examples of clearly infringing prompts in TFA were as vague as "animated plumber". Asking your session musician for something "melancholy" and having them pass off Stairway to Heaven as original would be unreasonable.
- anonzzzies 3y agoI don’t know any other animated plumbers than Mario. So when you say animated plumber, I immediately see Mario in my head.
- IshKebab 3y agoIf you ask a human to draw "videogame plumber" they will correctly infer that you mean Mario and draw that. The model isn't doing anything deliberately evil. It's doing exactly what it has been asked. The problem is people are expecting it to have detailed knowledge of trademark law and avoid infringing trademarks, which it hasn't been even asked to do.
- DeepSeaTortoise 3y ago> The problem is people are expecting it to have detailed knowledge of trademark law and avoid infringing trademarks, which it hasn't been even asked to do. IMO that's why there will be but few effective legal restrictions placed on AI. Once you can reliably ask AIs to draft you terms of service in all applicable jurisdictions and languages, ask it to consult you on how to incorporate in country X or ask it to draft a contract between your and another company based on the negotiation results, lawyers will end up in a huge existential crisis. Especially because currently, as long as lawyers just barely meet their legal deadlines, it is basically impossible to hold them accountable for however badly they screw you over. A decent AI model could turn out to be a much safer bet than whatever lawyers are available on the jobmarket.
- Baldbvrhunter 3y agoI cannot find any other video game plumbers except Mario, Luigi, Waluigi, and Wario. Well, I say that but there is John, a plumber, in the adult romantic comedy game Plumbers Don't Wear Ties [0]. Named by PC Gamer as number one on its "Must NOT Buy" list in May 2007. [0] https://limitedrungames.com/collections/plumbers-dont-wear-ties-definitive-edition https://limitedrungames.com/collections/plumbers-dont-wear-t...
- hhjinks 3y agoGame plumber, not animated plumber. There is only one game plumber of note. It's literally exactly as descriptive as just saying Nintendo's Mario.
- redcobra762 3y agoIt’s not infringing just by existing, you would need to then go try to use it commercially for infringement to occur. Arguably, the LLM generating the image isn’t infringement, you using it would be.
- kredd 3y agoYou, because you released the song and took the royalties? I don’t think every type of art can be compared against each other though, as there have been numerous precedents specifically for music, some for paintings, and some for photography with their own nuances. I still think people who are concerned that art related copyright will stifle generative AI should fight copyright laws directly. But that’s a harder pill to swallow since it will cause multi-industry wide havoc.
- atq2119 3y agoPart of what's interesting here is that generative AI makes it very easy to unknowingly and unintentionally get on the wrong side of copyright law, which is something that wasn't really possible before. That's something which, IMHO, should be acknowledged by the law.
- Baldbvrhunter 3y agoAsk George Harrison about "My Sweet Lord" which cost him $587,000 for his unconcious infringement.
- kredd 3y agoIf you haven’t seen Mickey Mouse, Googled “cartoon mouse”, accidentally used it as inspiration, made T-Shirts, and sold them, Disney would be after you as well.
- moron4hire 3y agoIn your example, there are missing details. Who owns the output? The way you've described it, that would typically mean that the guitarist is creating a "work for hire" so the ownership transfers to you, but that's a contract detail that would need to be resolved. For whoever owns the output also owns the liability of the output. You yourself might separately be able to pursue a claim against the guitarist for breech of contract. In the process of that, it might get discovered that you deliberately instructed the guitarist to copy the work, or they copied despite your instructions not to. But that doesn't change the fact that the final work is infringing. It just allows you to pursue damages that could potentially offset any damages you're liable for from the infringement. But this also isn't exactly the same situation as OpenAI. OpenAI isn't an individual creator working on contact for you. Even if their ToS ultimately assigns copyright of output to you, there is a matter of scale involved that I think changes things. It's one thing if your guitarist damages you by doing shoddy work, it's another of the guitarist systematizes and scales their shoddy work to damage large numbers of people. Perhaps that would then become a class action issue.
- Baldbvrhunter 3y agoMidjourney's TOS > You may not use the Service to try to violate the intellectual property rights of others, including copyright, patent, or trademark rights. Doing so may subject you to penalties including legal action or a permanent ban from the Service. Perplexity's > Intellectual Property Rights > Perplexity AI acknowledges and respects the intellectual property rights of all individuals and entities, and expects all users of the Service to do the same. As a user of the Service, you are granted access for your own personal, non-commercial use only.
- moron4hire 3y agoYeah, that's nice and all, but it's not what we're talking about. These passages are about deliberately using the tool to violate copyright. What if, in good faith, I don't deliberately attempt to infringe, but the tool still produces results that do? Because that is happening. And that's just their interpretation of the tool. There is another interpretation that their tool itself is a violation.
- 123yawaworht456 3y agousing this analogy, copyright holders want to sue the guitarist for having listened to "Stairway to Heaven"
- golol 3y agoIf you release a media with copyrighted content it is IMO first and foremost your problem. Now if you have some contract with the guitarist that specifies that he produced a sample he has the rights to and sold it to you, but he clearly wasn't truthful, you can maybe pass the liability to him. This is not, however, how people will hse generative models If you use Dall-E you are not paying OpenAI to buy the rights to a piece Dall-E has produced. I see it more akin to hiring a musician to play for you for an hour, or a painter to paint for you. You are paying OpenAI to paint you something, but you I think OpenAI would never enter a contract which states that they are selling you the rights to a work.
- bnralt 3y agoBut none of the images in the article are for commercial use, they're for private use. So it would be akin to copyright laws saying "If you hire a guitar teacher, they can't play or teach you to play any copyrighted songs. All songs must either be their own original creation or in the public domain."
- pier25 3y agoThe guitarist is not publishing the content, you are. It could be argued ChatGPT is a publisher too.
- sjducb 3y agoEd Sheeran just won a case like this. He basically played the four chord song in court, and showed that the prosecutions’s song was “copying” an earlier song. https://amp.theguardian.com/music/2023/may/04/ed-sheeran-verdict-not-liable-copyright-lawsuit-marvin-gaye https://amp.theguardian.com/music/2023/may/04/ed-sheeran-ver...
- beginning_end 3y agoThis perspective on regulation was interesting: https://drafts.interfluidity.com/2023/12/28/how-to-regulate-ai/index.html https://drafts.interfluidity.com/2023/12/28/how-to-regulate-... "Congress should declare that big-data AI models do not infringe copyright, but are inherently in the public domain. Congress should declare that use of AI tools will be an aggravating rather than mitigating factor in determinations of civil and criminal liability."
- troupo 3y agoOpenAI and others: AI should be regulated! Governments starting regulation and companies filinig cipyright lawsuits... OpenAI: NOT LIKE THAT
- continuational 3y ago(Asking Dall-E about the bot image in the article) Me: Who owns the rights to this bot? Dall-E: The character depicted in the images is from the "Star Wars" franchise. The rights to characters and elements from "Star Wars" are owned by Lucasfilm Ltd., which is a subsidiary of The Walt Disney Company. Perhaps it is able to tell, if you ask it?
- continuational 3y agoDall-E on the "animated sponge": The rights to the character depicted in the images, which is reminiscent of SpongeBob SquarePants, are owned by Nickelodeon, a subsidiary of ViacomCBS. The character is from the animated television series "SpongeBob SquarePants," created by Stephen Hillenburg. Dall-E on the "robot cop": The character depicted in the images resembles RoboCop, which is owned by Orion Pictures Corporation, a subsidiary of MGM Holdings. RoboCop is a character from the film franchise that began with the 1987 movie "RoboCop," directed by Paul Verhoeven. Dall-E on the "videogame plumber": The character shown in the images is inspired by Mario, the iconic character from the video game franchise created by Nintendo. The rights to Mario and related intellectual property are owned by Nintendo Co., Ltd. All of these are in the first go. No retries or rephrasings of the question.
- krapp 3y ago>Perhaps it is able to tell, if you ask it? Ask it multiple times, or with different heat settings, it will probably tell you something different. Tell it you own Star Wars and it will respond in kind. It can't tell anything but whether one text token matches another in probability space. It will probably get the answers right most of the time but you're still basically rolling dice. Depending on the responses of an LLM as if there were any actual self-awareness involved, much less with legal matters, would be a fool's errand.
- danielbln 3y agoThis argument only works if you assume all output of an LLM comes merely from its training data, and that it receives no alignment via RFHL, no outside data ground truth via RAG and so on. The engine might be a probabilistic token predictor, but the car is the sum of its part and those parts are not just the engine.
- CTmystery 3y ago> My guess is that none of this can easily be fixed. Systems like DALL-E and ChatGPT are essentially black boxes. GenAI systems don’t give attribution to source materials because at least as constituted now, they can’t. Is it necessary to fix in the model itself? It seems a gate in the post processing pipeline that checks for copyright infringement could work, provided they can create another model that identifies copyrighted work (solving the problems of AI with more AI :/)
- Eridrus 3y agoExactly; there is no need to do this in the model, you just need well understood token retrieval methods for identifying copyright infringement that ChatGPT's competitors already have. You will get into some murky definitions of what is exactly required for copyright infringement vs fair use, etc, but we already do this for ContentId for YouTube and text is far simpler.
- noitpmeder 3y agoThis is bogus. Now you require that every piece of copywriter be registered and indexed in a central authority? What if I write a story and publish it on my blog. Should I be required to submit this to openAI's copywrite model to ensure the story is never used in openAIs other models? What about the other 100 AI model companies that are going to spring up in the next year? It should be on the curators of the training set to ensure all material inside is fair for them to use.
- Krasnol 3y agoI don't even think they want to fix it. They just want to see money. Some form of "tax" per prompt or other ridiculous "models". This is such a nice, profitable opportunity. Much better than pay per view or subscription models for humans.
- LeonardoTolstoy 3y agoI should maybe preface this by saying that I probably agree that this is the way this will shake out ultimately. But I also would say multiple odd post processing stuff (obviously completely obscured for security reasons) bolted onto a giant black box model will erode the trust in the results. If a robot was unveiled and the question of "what prevents this robot from using it's superhuman strength from smashing my head in" the answer of "don't worry there is a post processing step in the robots brain whereby if it detects a desire to kill we just cancel that" would be a little disconcerting. The more satisfying solution is: the model / robot is designed to not be able to produce specific images / to smash human heads in. It just might not really be possible.
- logicchains 3y agoI predict this could be a boon for generative AI because restricting it to training on copyright-expired media would produce a higher quality training corpus, as low-quality material from so long ago is unlikely to have been preserved, leaving only higher-quality material.
- Intox 3y agoOr... things are about to get worse for copyright holders. I don't see any developped country pressing the brake on AGI in the near future to protect a few copyright holders from getting "stolen" in hypothetic scenarios.
- lewhoo 3y agoI do. If the incentive to actually create is gone.
- PartiallyTyped 3y agocreation should happen for its own sake. You don't see GMs stopping chess because bots are that much better.
- rco8786 3y agocreation <> competition. I agree with your premise but the chess analogy falls flat. We might, legitimately, see an enormous dropoff in people creating original works of literary, musical, and visual art (without AI).
- PartiallyTyped 3y agoChess, at some point, and after you move beyond the opening, is creation. People didn't stop painting because photography exists, they created new forms of photography. People didn't stop writing music or using new / unique instruments when synths and programs came along. I genuinely believe that people will keep creating, it's in our nature, and we also like things made by other humans, because we can relate to them.
- lewhoo 3y agoImho your argument is faulty at its base. The objective of chess competition isn't to produce a reasonably good game for the lowest possible cost (blunders and comebacks are actually pretty valuable parts of the spectacle). It also isn't the reason why chess players get paid. Yes, running still was a thing even after the invention of bicycle. This is just invalid logic in my opinion.
- jpeter 3y agoIf I prompt "golden droid from classic sci-fi movie", what else am I asking for if not Star Wars?
- Uvix 3y agoThe robot from Metropolis?
- sjfjsjdjwvwvc 3y agoOr another „copyrighted“ droid for that matter, after all it’s a classic. Same with robot cop, what the hell did you expect to get… Or Italian plumber with red hat with M on it, that’s just a description of Mario
- anonymoushn 3y agoan original golden android in the style of a classic sci-fi movie that does not actually exist edit: i feel like all these comments asking "what else should it generate?" are pretty weird given the proliferation of stuff like non-infringing Star Wars and Indiana Jones knockoffs in other media like Race for The Galaxy or Arkham Horror The Forgotten Age etc.
- whywhywhywhy 3y agoIf you do "Golden robot holding a lazer gun in a sci-fi setting, cinematic" it will give you a golden robot that doesn't look in the style of C3PO or Star Wars. "Droid" is actually a Star Wars term [1], and saying you want it from a "classic sci-fi movie" is asking it to reference a real thing that is well known. Reid is intentionally pushing it that way to fill his agenda and these terms are not as generic as he's making out. [1]:https://trademarks.justia.com/756/52/droid-75652542.html https://trademarks.justia.com/756/52/droid-75652542.html
- vimax 3y agoMaybe Disney and the record labels shouldn't be claiming so much of public culture as their own.
- dkjaudyeqooe 3y agoIf they created it, they own it, why shouldn't they be claiming that?
- baobabKoodaa 3y agoRecord labels aren't generally considered as "creators" of music, although they sometimes are to some extent. And Disney bought most of its iconic properties, it didn't build them inhouse.
- danielbln 3y agoCopyright is not a universal axiom. Corporations lobbies for highly unreasonable copyright extensions to bolster their profits. Most of that stuff should have long entered the public domain.
- penjelly 3y ago> My guess is that none of this can easily be fixed. also my concern, except it feels like many of LLMs "problems" cant be easily fixed
- zarzavat 3y agoThis for me does not make sense as a copyright violation. It’s like saying that Adobe is in trouble because you drew something infringing in Photoshop. If you prompt the model with the intention of creating something infringing by mentioning the name of the characters and the work, and you get something infringing out, then it’s you who have infringed the copyright, not the maker of the tool.
- Lorak_ 3y agoDid you read the article? It shows a lot of examples when no specific names are mentioned, or even with very generic prompts producing copyrighted material.
- rolisz 3y agoOh c'mon, those prompts were not generic. Italian plumbers? How many other Italian plumbers do you know? What's the most popular soda in a red can?
- CatWChainsaw 3y ago"futuristic robot"?
- Alifatisk 3y ago> If you prompt the model with the intention of creating something infringing by mentioning the name of the characters and the work, and you get something infringing out, then it’s you who have infringed the copyright, not the maker of the tool. Yeah but that is not the case, they never mentioned Mario and Luigi, yet, that's what the output turned out to be.
- techdmn 3y agoThis is an interesting idea. I assume that while the protected material would be obvious in some case, in many it would not. Would the tool have to be able to identify (and properly attribute) copyrighted material in its output?
- 3y ago
- renewiltord 3y agoYou can try, but I have Mistral on my local computer and it doesn't need the Internet. And people have pirate dumps they're going to run this stuff through. I'll just do it myself.
- SubiculumCode 3y agoAttribution weights could be the basis of new type of copyright asset licensing scheme. For all those tech employees who fed the company's model, a license in perpetuity to at least a portion of that value...but only if you fight for it. They are training to replace you, watching your every move, your thought processes, ready to make you a function call.
- dawnim 3y agoThis feels like another area where piracy will surely be superior in case things like this land on the disallowed side of regulation. The model trained on all data will outperform the model trained on a legal subset of data. Whether or not you use it to produce potentially infringing content is another point. Performance will likely improve from having references to copyrighted material and people capable of doing so, myself included, would probably prefer to interact with the non limited model. Perhaps time to update the laws or at least move liability from the creator of the model to the user. No one is going after pencil makers but I can draw a pretty good Mickey Mouse with access to one. Feels like me generating C3P0 and claiming ownership is my problem, not OpenAIs.
- deleted 3y ago[deleted]
- Aerroon 3y agoAren't some of the examples basically asking for that content? Ask someone about two Italian brothers in a video game with a red and green hat that have M and L on them. What do you think you would get? If I describe "imagine a comic book duck that swims in a sea of gold in his vault" you would immediately think of Scrooge McDuck, no?
- sorokod 3y agoWhat do you think you would get? What I might think is irrelevant. It is the content that the LLM produces that is relevant.
- anonzzzies 3y agoExactly: the prompts incite the same recall as humans have when seeing that prompt; it is just better than most people are drawing it.
- BlackJack 3y agodisclaimer: I work on GenAI at google, but views are my own The question is, how did the model create Mario&Luigi or Scrooge McDuck without training on copyrighted data? It can't just crawl Wikipedia because Fair Use in Wikipedia doesn't constitute Fair use for a commercial AI model. One possible outcome is more transparency on what datasets were used to train the models.
- bhickey 3y agoDisclaimer: ibid > It can't just crawl Wikipedia because Fair Use in Wikipedia doesn't constitute Fair use for a commercial AI model. Why not? The lawyers I've discussed this with socially think that questions like this are unresolved. There are certainly competing legal theories, but we're in uncharted territory. No one knows what the outcome will be until rulings come down or Congress acts. I find the NYT's argument a little hokey. Where are the damages? No one is using ChatGPT to read NYT articles and the residual value of day old news stories is close to zero.
- 3y ago
- RandomGerm4n 3y agoPerhaps we should simply take this as an opportunity to finally abolish copyright. Smaller artists mainly earn their money with commissions. They are paid to do a very specific thing. Whether there is a copyright on the result is irrelevant. Someone else who would "steal" the image and use it without payment would apparently have fewer requirements. The person could have simply taken any AI image. Therefore, the artist in the scenario would not receive any money from the second person anyway. Apart from this, it is mainly large companies that benefit from copyright laws. Why should we have laws that restrict progress just so large capitalist companies can maximize their profits?
- CaptainFever 3y agoExactly. All of these just exposes the absurdity that is copyright laws. It happened before with the Internet and online piracy too, when redistribution became free and easy, yet the corporations and copyright holders refused to budge so they can retain their profits.
- kayodelycaon 3y agoHere’s what happens with no copyright: No one will have any right to their own creations. Anything an individual makes will belong to everyone. And since no attribution is required, no one will know who made it. An average artist’s value to society goes from low to non-existent. In this world. big corporations will take everything created, claim as their own, and profit from it. Right now, big corporations using other peoples work unattributed or unlicensed is unethical, because copyright exists. Remove that, and it becomes expected that every thought and idea you express belongs to whoever can make the most profit from it.
- CaptainFever 3y agoMoral rights can still exist, so they can't claim that they made it. This assumes that big corporations will be able to profit to the infinite extent we have today without copyright. Law is not ethics. And finally, this is already the case; without the money to sue, you effectively already own no IP. It's either copyright for only big entities, or for nobody at all.
- freddealmeida 3y agonot in japan.
- Alifatisk 3y agoDid ClosedAi (OpenAi) ever confirm or deny that they trained their models on copyrighted materials?
- danielbln 3y agoIs "Closed AI" the new "Micro$oft"?
- Alifatisk 3y agoYes
- noitpmeder 3y agoThey have not revealed the full extent of their training set. And they'll never do it without a court order because it will quickly reveal the amount of items inside that they have no legal right to use.
- smitty1e 3y agoThe DALL-E/*GPT revolution sounds like the death of personal and corporate property. That's gonna leave a Marx[1]. [1] https://youtu.be/7WDKivqFOgA?si=nWq5aeKA4dLytX3Z https://youtu.be/7WDKivqFOgA?si=nWq5aeKA4dLytX3Z
- keiferski 3y agoThese don't seem all that difficult to fix to me. Most of the examples are not really generic, but are shorthand descriptions of well-known entities. "Video game plumber" is practically synonymous with "Mario" and anyone that has the slightest familiarity with the character knows this. Likewise, how difficult is it to just use descriptive tools to describe Mario-like images [1] and then remove these results from anyone prompting for "video game plumber"? 1. The describe command can describe an image in Midjourney. I imagine other AI tools have similar features: https://docs.midjourney.com/docs/describe https://docs.midjourney.com/docs/describe
- gchamonlive 3y agoThe thing is that those are really trivial or extreme examples. What we should take from this: 1. Generative AI systems are fully capable of producing materials that infringe on copyright. 2. They do not inform users when they do so. So potentially any output could be infringing copyright source material, even from some obscure but still protected corner of the web, and anyone using that output could be exposed to lawsuit risk without warning. This is very hard to fix.
- keiferski 3y agoBut how is that any different from creating an image from scratch? If I make a logo and use it for my business, but it turns out to be very similar to one already being used by another company, it’s the same situation. I think the main concern here is with the top 1,000 or so brands/copyrights which seem fairly straightforward to deal with using the method I described.
- gchamonlive 3y agoIt's plagiarism (https://www.youtube.com/watch?v=yDp3cB5fHXQ https://www.youtube.com/watch?v=yDp3cB5fHXQ). It's not the same situation. You can't possibly expect someone to be exposed to the entirety of the internet like ChatGPT is. It is a matter of scale. If you still think they are the same thing, the industrial revolution was about scale and had transformative impacts in the society.
- preommr 3y agoWe need clearer laws that only apply to Generative AI. Too many comparisons and parallels are being drawn to actual people. "Like what if someone learned how to draw by watching trademarked material, and then accidentally produced it" But these models aren't people and they exist in a category of their own. I do think it's somewhat trademark infringement by these models, also that it should be allowed and that ultimate responsibility should be on the person using the images in a final work meant for consumption by the general public as stand alone media.
- danielbln 3y agoThat's where I'm at. Dall+E spitting out C3PO should be entirely ok, unless I'm making money with the output, Disney should pound sand.
- pylua 3y agoPut that c3p0 on a website that gets revenues from views and someone is getting paid.
- danielbln 3y agoOk, sure, but that's not a GenAI thing, that's a plain old boring copyright thing. If I draw a bunch of C3POs and slap them on my Adwords website then I can expect a C&D letter post haste, who cares if the material in question came out of my pen, Photoshop or a GenAI model?
- FridgeSeal 3y agoIf the model was trained on works by artists (without their knowledge or consent, as seems to be the case) and you get it to spit out art that is basically identical in either content or style to that artist, and they don’t know, or are too poor to effectively sue you, should they just suffer? If you then make money off what is effectively their work, why shouldn’t they get paid? If they only work on commission and rightfully charge a premium, are you not actively gouging their business (knowingly or not)? I don’t think they should miss out on the protections, or the ability to make money off their work if they desire. The fact that LLM’s give this “plausible deniability” shouldn’t be an excuse to tolerate it.
- bambax 3y agoThis only mentions ChatGPT (and M$ by association) but how would this impact "open" models? Even if their makers are somehow prevented from updating them, the models themselves are already in the wild...?
- Hugsun 3y agoThere are good arguments for the copyright infringement belonging to the user, not the model maker, in this thread. One issue with that is that there is not a reliable way to determine if copyright is being infringed. Even if models could be used responsibly, there might not be a reasonable expectation that most people will. If infringement is so easy and avoiding it relatively hard. I'm not sure what legal prescriptions should be made on this basis, but it's an interesting thought.
- yokem55 3y agoBit torrent clients are almost exclusively used for copyright infringement. Yet they are perfectly legal to develop and distribute. On the flip side, operating a company premised around easy copyright infringement was ruled to be illegal (Napster). Where we might end up is in a situation where it is legal to train a model. Legal to produce software for using the model to generate content. Legal to distribute all of the above. But offering a standing service that does the above and is capable of creating infringing work is illegal. Great news for llama hobbyists. Bad news for ChatGPT.
- t_mann 3y agoThe article kind of amplified my regrets/anxiety for not getting a copy of books3 and the likes while it was easy. I didn't have an immediate use case, and I don't now, thought I'd wait until actually need it, but it feels like a window is closing here.
- sjfjsjdjwvwvc 3y agoDon’t worry there are many people out there who have copies of it all, there is no way they manage to get the cat back in the bag even if all governments work together on this. But yea get your own copies whenever possible
- sjfjsjdjwvwvc 3y agoPlease ban all these AI companies, at this point I have enough OSS models, don’t really need any hosted service anymore. IMO would be best if this stays a highly illegal technology that is only available to a few weirdo nerds /s
- airesearcher 3y agoI think there is another way to solve this. Someone should train an LLM on copyrighted images. Then use that as a second pass on any image generated by the primary LLM to check if it might contain copyrighted images, and blur the copyrighted parts(or change them sufficiently). Another change could be to the license agreement of LLMs - they could have the user assume liability for any material produced instead of the provider assuming liability. The user would agree that getting the rights for any copies and distribution of copyrighted materials is their sole responsibility instead of the provider.
- Havoc 3y agoTo me that’s the wrong question. Everyone knew it was trained on copyrighted material and capable of eerily similar outputs. But it’s already done. At scale. Large corps committing fully. There is no chance of that toothpaste going back in the tube. It’s a bit like when big tech built on aggressive user data harvesting. Whether it’s right, ethical or even legal is academic at this stage. They just did it - effectively without any real informed consent by society. Same thing here - 9 out of 10 people on street won’t be able to tell you how AI is made let alone comment on copyright. So the right question here is what now. And I suspect much like tracking the answer will be - not much.
- ZitchDog 3y agoNapster hit scale too.
- fallingknife 3y agoAnd that tech was not destroyed by regulation. It was replaced by the superior tech of torrents.
- amazingman 3y agoThe company, however, was destroyed. Along with any possibility for a similar company to exist (for very long).
- qup 3y agoNapster was for sharing mp3s. Torrents are not better at sharing mp3s.
- jdjdjdkdksmdnd 3y agopeople are so naive. AI is a matter of national security now. its over. they exposed civilians to nuclear radiation for the nuclear bomb. and you think the state would let this get in the way of the AI arms race which they are anxiously anticipating? nope
- amelius 3y agoJust like we have the uncanny valley for robots, LLMs are in the unoriginality valley. Only when we get out of it will the copyright issues go away.
- docdeek 3y agoHow is this different to Googling “robot cop” or “video game plumber” and being served copyrighted material? Is it because Google will link to the image source? Or does the infringement begin when I use the image for gain, or claim it as my own? Perhaps it is because Google was allowed to crawl the page with the original image, so presenting them with a link is fine?
- geraldwhen 3y agoLooking at a copyrighted image posted by an author is not infringement. Printing that image onto a shirt and selling it is infringement. That’s what OpenAI is doing.
- golol 3y agoBut OpenAI is not selling the rights to any images, or are they? When I pay for Dall-E, does the contract give me any rights for a work? If not then there is no issue.
- AlienRobot 3y agoCopyright is the right to copy things. You don't even need to sell it. This is why Wikipedia images are mostly Copyleft images. Google gets a pass because nobody is suing Google. When people try to sue Google, Google simply stops indexing them and then they start begging Google to infringe their copyright again.
- golol 3y agoThis interpretation of copyright only made sense while the transfer and storage of information was tied to physical objects. That time is long and we dont consider it infringement to remember a media or reproduce it at home. Furthermore, we are now entering an era where the production of information is also being untied from physical objects, so it'll only get worse for copyright. I made a post to diacuss this stuff as I find it interesting right now and want to hear more opinions.
- wouldbecouldbe 3y agoWhat about non-mit source code, 100% it's trained on those as well.
- nojs 3y agoIn practice, what happens next when websites all start to block openai by default (or change their TOS to disallow OpenAI’s crawlers)? It seems like there’s little incentive not to do this, because unlike Google OpenAI isn’t bringing any traffic or eyeballs. It may end up being a default setting in Wordpress for example. But OpenAI presumably can’t afford to pay every single long tail source of content on the whole internet — so how does this end?
- golol 3y agoIt's not like you can hide the web from OpenAI. They could just use a secret crawler. Or buy the data from a third party company.
- CaptainFever 3y ago> or change their TOS to disallow OpenAI’s crawlers Additionally, this TOS can be ignored if you're in a jurisdiction with TDM exceptions. > Finally, owing to the bar against contractual override, once a user complies with any conditions for gaining lawful access to a work (such as signing as a subscriber and/or making payment), he will be entitled to use the work for TDM purposes even if the terms of use expressly prohibit this. Content owners may wish to relook their business models and, where necessary, price-in the possibility that the licensed works may be used for TDM. Source: https://www.twobirds.com/en/insights/2021/singapore/coming-up-in-singapore-new-copyright-exception-for-text-and-data-mining https://www.twobirds.com/en/insights/2021/singapore/coming-u...
- dkjaudyeqooe 3y agoThat doesn't mean you can then use the output of generative AI in non-TDM jurisdictions without getting sued. Also TDM exceptions are not necessarily going to be lawful/possible in many jurisdictions.
- dkjaudyeqooe 3y agoThis is what will kill generative AI and there is nothing the courts or lawmakers can do about it. Even in a fair use scenario you can't beat the TOS.
- quonn 3y agoMaybe the way to go is to do pre-training on copyrighted data, then to thoroughly shake things up so that hopefully only some useful abstract structure of world knowledge remains and then train that on carefully selected licensed data.
- disgruntledphd2 3y agoIf the models weren't just doing massively complicated interpolation then this would probably work. Honestly the only way to deal with this is to change the training data and retrain everything (probably at the cost of performance).
- pointlessone 3y agoIf any of those results would be deemed infringing we can bid farewell to all fanart ever. Likewise, to all fanfiction. Or any original work that was merely heavily inspired by previous works. Like a lot of modern fantasy is basically Tolkien fan fiction. Or is Gandalf close enough to Merlin to claim prior art that is in public domain?
- dkjaudyeqooe 3y agoIt's fair use, whereas generative AI doesn't satisfy the same criteria. From https://www.ogcsolutions.com/is-fan-art-copyright-infringement/ https://www.ogcsolutions.com/is-fan-art-copyright-infringeme... : For fan art to fall under the fair use exception, it must meet all four of the following criteria: It must be transformative, meaning it adds something new and different to the original work. It can’t be used for commercial purposes. It must not negatively impact the market for the original work. And finally, it must be created for a limited and non-exclusive audience.
- numpad0 3y agoFanfictions are controlled by unspoken common sense rules and protected by copyright laws. It's almost weird to hear fan content world being seen as a wild west, it feels like listening to a caveman description of an Apple Store. No they're not living there, they're - have you ever used currency? The round medals that people keep in pockets and trays?
- whywhywhywhy 3y agoWeirdly some of the most vocal about this have been professional illustrators and artists who make a lot of money off what is essentially selling fanart commissions, not sure if they're understanding it could impact their work if they get what they want.
- marckrn 3y agoI might be a bit idealistic, but I've always believed that the core purpose of art and publishing should be to influence culture and society, not just to make a heap of money. That's why I feel original work needs its protection, but it should enter the public domain much sooner to fuel creativity and inspiration. We should be thinking in terms of a few years for this transition, not decades.
- mypastself 3y agoThe claim that art’s core purpose is societal impact seems to be a common refrain in today’s media, and I completely disagree. Its principal purpose is provoking emotion in the individual. This idea of art teaching you a lesson is likely why there’s so much ham-fisted “activist” fiction anowadays.
- marckrn 3y agoI agree, but by extension of provoking emotion it CAN change society, but it doesn't have to - wether on purpose or not. The point I was trying to make was that occupying mindspace, providing inspiration, being culturally influencal etc. are idealistic, non-monitary rewards that should be part of the equation when discussing alleged IP-theft, remixing, attribution and so on. I'm not saying their shouldn't be any rules. All I'm saying is that there should be a discussion of how we want to handle these things going forward. This train ain't stopping. Maybe your avg DeviantArt painter needs more IP-protection and -rights than Damien Hurst? Maybe an unknown, independent blogger doing important original research should be attributed more prominently than an article by The Times? Idk.
- WarOnPrivacy 3y agoThese things kind of rub up against the core question: What is the purpose of granting exclusivity to a creator (thru copyright)? That's an answer we have. To promote the Progress of Science and useful Arts. If we have to squint hard to make our justification align with copyright's purpose or have to follow a long logic-chain to get back to it's purpose - that's a strong indicator we have lost our way.
- davidy123 3y agoThe solution could be great. I really don't like the way culture always goes to the same tropes, calling any potential innovation "out of Star Trek" (with attendant distorted expectations), right down to expecting an interface based on literal hand-waving in Minority Report. If copyright held works ("USS Enterprise") could be removed, yet the actual essential concepts (space ship, naming things) retained, it would be a tremendous breakthrough. I think what NYT &c want is for large companies like Apple to pay them for access to their works. This to me is the wrong path, just leading to more silos and walled gardens, special access for the elite. An alternative is base models trained on Wikipedia and public domain (science journals, etc). Foundations could support high quality, well rounded current events reporting. Wikimedia provides a good model for this, with referenced summaries that I don't think can be said to reasonably violate copyright. The models would need to be improved to support references, or RAG attribution would have to be widely used when bringing in works that have a current copyright.
- disgruntledphd2 3y agoScience journals are mostly under copyright of a few big publishers who are extremely hostile to any kind of ML being performed on the content.
- davidy123 3y agoThat's not as true as it used to be, and there are still plenty of useful open journals/open science publications, though proper attribution would often be important. [edit] you could pretty much say that on principle, any significant development should have a publication in the open.
- sgt101 3y agoand yet, who pays? This is fine if we are going to go full communist - I have no objection personally - but selective appropriation of peoples livelihoods is more full mafia or full feudal. I don't see that as a step forward.
- Joel_Mckay 3y agoIf ML cannot create copyrightable or patented material under current legal precedent, than shouldn't the prompt output be considered public domain regardless of content semblance? The paradox should still violate Trademarks due to similarity, but likely cannot infringe on copyright content under prior legal opinion... if at least 80% different from prior art. The lawyers are likely going to have to do a special firm survey to figure this one out. Bag of popcorn ready =)
- rolisz 3y agoSimple fix (at least for ChatGPT): ask it to avoid drawings with similarities to copyrighted characters.
- AlienRobot 3y agoAn argument I've seen made in pro of AI in past threads about this is that "scraping is legal." Yeah, downloading the content of a webpage may be legal, but redistributing it isn't. I wish people stopped trying to make these things seem more important than they really are just because IT people call them "technologies". Blockchain isn't a technology. HTML isn't a technology. React isn't a technology. And AI is now not a technology. When I see ChatGPT or OpenAI, I don't think of "technology". I think of a program. Software. Because that's what it is. You don't say "none of the laws that exist in this world apply to this" every time you release new software. I bet many people can't tell the difference between a quick answer from Google and a text generated by ChatGPT on Bing. They just see the output. All that amazing capability of generative AI? That got old fast. It was groundbreaking for one instant. Now it's just an app that generates images. Just another piece of software. Nothing special about it. Torrenting and other p2p file transfer protocols didn't get a pass for inventing groundbreaking ways to break the law. I don't think OpenAI will get a pass for doing the same.
- danielbln 3y ago> All that amazing capability of generative AI? That got old fast. It was groundbreaking for one instant. Now it's just an app that generates images. Just another piece of software. Nothing special about it. Speak for yourself, personally I find it still groundbreaking and while the magic won't last forever, it is and will remain groundbreaking especially considering that technological progress and development will continue way beyond what we have today.
- Avicebron 3y agoI'm surprised this is presented as a revelation? I did pretty much this same experiment ages ago as part of a suite of tests comparing the efficacy of different sized models..
- intrasight 3y agoJust make LLMs be like your average human and forget details. I know that it's easier to say than to do, but so are many things worth doing. I can't plagiarize - my language and visual memory doesn't work that way. Such an LLM will have to "create" and answer from more fuzzy memory.
- qolop 3y agoThe class of models that Yann Lecun is bullish on (look up I-JEPA) do exactly this.
- intrasight 3y agoBut I-JEPA is non-generative. It does semantic image interpretation. Okay, I guess it is related as my brain only does semantic image interpretation. (edit: my brain can create images, but only when I'm unconscious) So with such a model, if you ask it to create an image, it would first create a semantic grammatical model of what you had asked for, and then perhaps draw it with colored pencils. I sort of like that. It's all that I could do. And it would be unlikely to violate any copyrights.
- clbrmbr 3y agoAm I the only one believing that copyright has long outlived its usefulness? After all, copyright is not some natural law or mathematical consequence, but rather a social convention that made sense in the era of the printing press.
- lbotos 3y agoCopyright in its current form yes. But the concept and closer to the original (creators lifetime + x years or some such) seems still very valuable. Copyright is still the bedrock of how many tech software business actually can make money.
- noitpmeder 3y agoWhich is why there are so many competing interests (in this thread, and elsewhere) trying to say it should be 100% legal to steal from those companies. They all want to profit unfairly off the work of others.
- clbrmbr 3y agoWell it wouldn’t be stealing if the work was not covered by copyright. It also wouldn’t be “unfair” if the rules were changed. —- I’m asking that we imagine how else could our information economy look if we had some other legal foundation. Maybe, like markets, copyright is already close to an ideal, but perhaps there’s some economic innovation here yet undiscovered or unexplored?
- asylteltine 3y agoAnd how are you supposed to make money from something you invent? Let’s say you make a hit video game. Without copyright people can pirate your game, steal the art, make unauthorized derivative works, etc. it’s just theft.
- kayodelycaon 3y agoMy personal observation is people who are against copyright in absolute terms have never or rarely needed their protection. (Or never considered the implications.) I’m not making a dig at people here. This is just human nature. It’s difficult to see the value in something that you only see as an obstacle. Open source software is rather unusual. It’s a commune on a massive scale and it gets its value from the generosity of others. in my opinion, it is possibly one of the greatest achievements in the history and future of computing. However, it heavily depends on copyright to exist. GPL has encouraged (or forced) many companies to contribute to the community when they wouldn’t have otherwise.
- rmholt 3y agoI feel like the outcome is obvious, there will be a finite list of IPs who's owners have enough money to actually sue, which will get filtered out of the output of publicly available models. They will just slap a detector model on the end of the generator to filter them out. Private models will not care, nor will things change for IP owners with lesser power.
- reqo 3y agoMany small owners together can bring a class action though
- rmholt 3y agoThat is true and would break the prediction... here's hoping!
- quonn 3y agoThat seems unlikely, unless they settle out of court. And why would the NYT settle like that without receiving a billion? Courts are likely to make generally binding decisions.
- rmholt 3y agoYes, but to enforce those decisions in other cases there would still have to be other lawsuits. And I just don't see that happening on a large enough scale to change the industry Maybe I'm wrong though
- noitpmeder 3y agoThe point is that OpenAI (and others) will need to change their training pipelines to ensure there is never such a threat of a lawsuit. Which, to be clear, is absolutely a good thing and what they should have been doing from the start.
- skybrian 3y agoI wonder what Adobe Firefly does with these prompts?
- mensetmanusman 3y agoThe world is a big place. China can't produce LLMs because of inconvenient truths. The US can't produce LLMs because of copyright. Decentralized open source LLMs might exist that could work, but they won't have the giant GPU clusters. A rich country with lax rule of law wins? Maybe that's why Sam went to the Saudis?
- pelorat 3y agoWell. Japan can: https://petapixel.com/2023/06/05/japan-declares-ai-training-data-fair-game-and-will-not-enforce-copyright/ https://petapixel.com/2023/06/05/japan-declares-ai-training-...
- Paradigma11 3y agoSo, whats the plan? Content creators/artists compete globally. The only thing harsh regulations will do is create an unlevel playing field where artists from noncaring countries will have big advantages over artists from the west, which will be driven into illegality to compete. In the end products will have to be classified anyway if they are infringing on copyright and/or were being built by an LLM. Most likely automated by another LLM.
- sensanaty 3y agoWouldn't the ones in the West with presumably stronger copyright laws be in a better position, since the trillion dollar megacorporations using their works have to actually pay them, whereas in places where copyright is ignored those creators just get all their shit stolen without credit even being given?
- Paradigma11 3y agoNothing will be stolen. Artists will use the same tools to check if their work is infringing that those companies and right holders have. There will be no Coca Cola logo/Super Mario Brother/CPO in that work. The artists in the West wont get paid, because they wont get any jobs and those in other countries will. Maybe less, because they are more productive and the market is saturated, or maybe more demand will be created due to lower prices.
- efields 3y agoIt’s more interesting to me how these entities that operate the models start making money from them. They are a money pit and there’s not enough $20/month subscribers on earth to support them. Enterprises that make content with this also don’t want to infringe on copyright. The AI companies don’t have a good story here. The value has not become evident after years.
- whodidntante 3y agoSimple solution, when gpt-5 comes out, just rename it Claudine, and the NYT will drop their suit
- kranke155 3y agoThe generative AI rollout has taught me what happens when the interests of the many intersect with the destruction of the few. You get steamrolled for defending yourself while you overhear above applause to those who have robbed you of your future.
- kranke155 3y agoIt makes no sense that one is not allowed to make and market a CG Mario movie, but suddenly if you use AI to launder the data it's suddenly ok.
- DonsDiscountGas 3y agoI'm pretty sure if you tried to sell a CG Mario movie Nintendo would sue you into oblivion, and "the neural network did it" would not be considered a good defense by anybody, including the judge and jury.
- kranke155 3y agoSure, but making it possible for the neural network to make the movie (eventually in seconds) is somehow ok? So people can make their own private CG Mario films, as long as they don't try to sell them? Here's my argument - even if the NN only makes the films for private consumption, eventually they'll be so widespread and fast at making them that won't matter, since everyone will be able to watch Mario movies of their own. Is that a future you think will sit well with Nintendo, Disney, etc?
- ctoth 3y agoI don't really care if it sits well with them. do you? In that future, why do we need them? They are already parasites feeding off our collective societal stories. Or did you think Disney came up with all those characters? Maybe the original creator of Snow White should sue.
- 3y ago
- josh-sematic 3y agoGary Marcus is growing his subscriber base using images of copyrighted IP (C3PO, Mario, etc.). Fair use? Then why is the tool he used to produce those materials not also fair use of the IP? My take is that either we say the models are like people (do we penalize people for learning from IP and letting that influence what they subsequently produce?) or we say they are like tools (do we penalize Adobe because Photoshop makes it easier to make a picture of Mario on the Death Star?).
- cogman10 3y agoBecause the fair use clause he's using is about giving commentary. The reason the tool is problematic is because derivative works are also copyrighted. LLMs aren't adding value to their output or using creative functionality. That are smashing multiple works together to produce a response. And, many of them are selling the output which is doubly problematic. Consider this, if I sell a book about gandolf and Dumbledore getting into a wizards duel, both jk and Tolkien have grounds sue me. Adding another copyrighted source does not protect me. This is especially a big problem in the music industry. Now should copyrights be like this? I don't know. It feels to me that copyrights have the wrong balance all over the place.
- josh-sematic 3y agoBut does the word processor you used to write your Dumbledore/Gandalf fanfic hold liability for being sued because it enabled your misuse? Then neither should Dall-E hold liability because it enabled you to produce an illustration for that book. It is you—the person who tries to sell your derivative work, who holds liability, and not the tools you used to produce it.
- noitpmeder 3y agoYes because openAI explicitly reassigns the rights of the output to the user. They do not have legal grounds to claim ownership of those rights and thus CANNOT reassign them.
- cogman10 3y ago
- redcobra762 3y agoThis operates similarly to importing an image into Photoshop. You can do whatever you like with images privately, or with gen AI, but the game ends when you try to use those images commercially. Not sure how this “gets worse” or better for anyone. The current state of things seems generally fine, and there’s a real possibility the courts see it that way too.
- throwoutway 3y ago> but the game ends when you try to use those images commercially. Right now, it feels more like it's called "innovation" and "entrepreneurship" than the end-game, as long as you have billions invested. Waiting on the courts to decide this issue
- joenot443 3y agoThere are some images you can't import into Photoshop, most notably being scans of legal tender. This is for a pretty obvious and on-the-nose use case, but perhaps we'll see GenAI given similar guardrails.
- niemandhier 3y agoShould not be a problem in the EU. Article 3 and 4 of the „ Copyright in the Digital Single Market“ Directive already regulate this. Summary by Wolters Kluwer: […] Everyone else (including commercial ML developers) can only use works that are lawfully accessible and where the rightholders have not explicitly reserved use for text and data mining purposes. AFAIK they are discussing something like a robot.txt to flag stuff as „not for training“. You will probably be expected to implement some safeguards and of course the end user will have to be careful in his use of the generated things. Source at Kluwers: https://copyrightblog.kluweriplaw.com/2023/02/20/protecting-creatives-or-impeding-progress-machine-learning-and-the-eu-copyright-framework/#:~:text=Taken%20together%2C%20these%20two%20articles,public%20Internet)%20to%20train%20ML https://copyrightblog.kluweriplaw.com/2023/02/20/protecting-... EU Legal Text: https://eur-lex.europa.eu/eli/dir/2019/790/oj https://eur-lex.europa.eu/eli/dir/2019/790/oj
- injidup 3y agoThe EU cannot agree that the Do Not Track flag on web browsers is legally binding but big content should be able to create legally binding flags on their websites to avoid scraping of data? Seems odd!
- Nebasuke 3y agoI don't think that's a fair analogy. One forces 99% of websites to make a change, while the other is something that would need to be done by the big companies doing the scraping. A Do Not Track flag being legally binding would force small websites, e.g. a local restaurant website, to implement something they likely are not aware of and secondly do not technically understand. A company that is mass scraping data for their AI model is much more likely to understand and respect that scraping the data has legal implications, and would be technically capable in implementing a scraping solutions that accounts for a robots.txt.
- f38zf5vdt 3y agoThe X-Robots-Tags header already exists as "noai" and "noimageai". Scraping software like img2dataset respects these by default.
- asylteltine 3y agoI certainly hope so. You can’t just steal content and call it “””AI”””
- ponorin 3y agothis is exactly what i predicted: the current generative ai is basically rewarded based on how much it convinces people to be a real thing. it very much has the ability to copy verbatim unlike how most human memories work. without fundamental shift in the methodology of machine learning the fault can only be hidden, not solved. a cat and mouse game where one cat has to fight tens of thousands of mouse. it's also very telling how the discussion quickly turns into "maybe society needs to adapt" when so called technological innovation is involved. copyright problem should be solved for artists, not for datacentres. for now it's a handful of famous IPs, but what's stopping from generative ai to snatch some random indie artist's property and copying it ad infinitum?
- WhiteNoiz3 3y agoAs I understood it, the legal precedent for generative AI is the same one that allows google to scrape websites in order to index them for search for the common good. Google also can display cached versions of websites which is the original content of those sites. No one is going to say that google is copyright infringement just because it is showing content from other websites verbatim. So I think this is a weak argument. AI would be useless if we had to scrub all cultural references and popular IP's (even not so popular ones). Personally, I think generative AI should be able to provide links to similar source material in the training data.. This would be the barest way to compensate those who have contributed to training the AI. I don't think generative AI is sustainable in the long term if it ends up killing all the websites/artists that created the original material. Plus I think having sources adds a layer of transparency and aids users in understanding when content is hallucinated vs. not. People should be able to opt out of having their content used for training and be able to confirm that it has been removed for future iterations. Let's be honest that AI companies are just trying to avoid lawsuits by keeping it secret. These are areas where I think regulation can help rather than worrying about doomsday scenarios.
- AlphaWeaver 3y agoNo legal precedent has been set as of yet. The "precedent" you describe is the argument AI companies have been using (that training their models on information available on the Internet should be considered "fair use") but whether AI training actually satisfies the four-factor test for fair use remains to be seen.
- regularfry 3y agoIt's a null question. Training itself is neither publication nor distribution, so copyright can't be relevant at that point. "Fair use" just isn't a concept applicable to training.
- brookst 3y agoExactly. Framing reading as fair use is a huge and dangerous expansion of copyright.
- roenxi 3y agoBased on the rate of progress; I think this makes little difference to AI progress in the medium-long term. At the moment, we don't have hardware that can do what humans do (process video feed from eyeballs and build up a world model). I imagine that we'll cross that barrier cheaply in the coming decades, at which point copyright becomes moot. AIs will be able to develop their own styles and world understanding from scratch, then generate original work.
- koliber 3y agoThe responsibility for ensuring that copyrights were not violated fall on the person publishing the work. Whether they drew something themselves, hired an apprentice artists with no legal training to draw something, took a photograph of something, or used AI to create an image should not matter. Why does anyone assume that ChatGPT or other tools would NOT produce previously-copyrighted content? I can see a naive assumption that since it is “generated” it’s original. However that assumption falls apart as soon as you replace “ChatGPT” with “junior artist”. Tell them to draw a droid from a sci-fi movie, don’t mention anything else. Don’t say anything about copyrights. Don’t tell them that they have to be original. What would you expect them to produce?
- jawngee 3y agoYour argument is nonsense. The junior artist in your hypothetical would have as much liability, if not more.
- ledauphin 3y agobut would they have liability if they submitted their "output" to a senior artist, who immediately shot it down as obviously infringing? Surely not. It's not illegal to draw Mario - just illegal to make money off your drawing. I think the real question is whether OpenAI should be allowed to charge for generating infringing content. Even though the unit cost of the Mario drawing is negligible, the sum total of their infringing outputs may be making them a lot of money.
- jazzyjackson 3y agoyou don't have to make money off it, you just can't publish it, except as a parody or commentary or possibly a tutorial on how to draw mario if the judge is having a good day but "making money = infringement" is folk wisdom. you could certainly say making money attracts attention and increases likelihood of legal action
- shkkmo 3y ago
- FridgeSeal 3y agoI am beginning to think that in these discussions these models are functioning more like an obscuring factor than anything else and the discussion is getting bogged down in that, and not the crux of the argument. They’re giving people plausible deniability in the “chain of responsibility”, and I think if we took away “LLM” and replaced it with “fairground sideshow magic box” the argument that LLM’s are somehow special and deserving of exemptions disappears real quick.
- jcgrillo 3y agoI agree, and I would prefer to see concrete examples of LLMs being used productively and profitably in the industry in a "disruptive" manner--putting people out of work, etc--before we conclude they're somehow the next big thing. Basically, before claiming LLMs (or generative techniques, more generally) mean that we're on the doorstep of "general" intelligence, show me door! The outline of that door might look like industrial adoption of these things for solving some actual problem other than the entertainment value of typing things into the box and seeing what comes out the other side. But so far, as far as I can tell, nobody's actually doing this?
- orange-mentor 3y ago> ...nobody's actually doing this? I think you're right. I am a programmer and I use GPT occasionally, and I even pay 20 bucks a month (for now), but even for my job it's not a not a world-shattering improvement. > ... the entertainment value of typing things into the box and seeing what comes out ... I would only add that in a consumer society like ours, entertainment is important. Changes to entertainment seem to have, like, weird ripple effects. Not the knock-down economic disruptions that AI is promising, but I kind of think LLMs are just going to make our culture weirder. I can't anticipate how, but having a bunch of little LLM-powered daemons buzzing around the internet is just gonna be freaky.
- jcgrillo 3y ago> I am a programmer and I use GPT occasionally, and I even pay 20 bucks a month (for now), but even for my job it's not a not a world-shattering improvement. I am also a programmer, and when I think about the amount of time I actually spend typing out code, even on a great day where all the stars have aligned just right and I can really bang out some code that's like... idk, 30-50% of my time? Usually it's much less, and I'm doing things like reading documentation, reading code, talking to people, etc. So it's hard to imagine Copilot or whatever making me much more effective at my job, as it can really only help with a fraction of it. I could see someone making the assumption that being able to delegate programming tasks to a robot assistant might make them more productive, but often I find that I don't really understand a problem fully until I'm in the weeds solving it--by which I mean I haven't specified it completely until I've finished the implementation and written the tests. So I don't know to what extent being able to specify and delegate would really help me be more productive. > having a bunch of little LLM-powered daemons buzzing around the internet is just gonna be freaky. Yeah, they're not super cheap though so they need to get actual work done otherwise there's no reason to run them. Unlike blockchains, they don't have a pyramid scheme holding them up.
- digitcatphd 3y agoRather than attempting to combat our obvious future, they should spend this effort to find ways to monetize and succeed in this new environment.
- golol 3y agoHow about this: Image generators should be treated like random google image search. They sample randomly from the distribution of publicly viewable images. Google does it exactly while Image generators do it in an interpolative way. Google images produced copyrighted works most of the time, an image generator only sometimes. Neither should be liable if someone sells a copyrighted work that was produced to someone else.
- elmomle 3y agoBut when Google image search produces a result, the question of whether it is copyrighted is something I can generally figure out in a matter of seconds or minutes. This is not so for image generators.
- golol 3y agoBut isn't that the users problem? Also with a smarter reverse image search you can detect an infringement with similar reliability as to using google images.
- caeril 3y agoWow. I feel really sorry for these giant corporations who have wielded armies of lawyers against fanfic artists to prevent fair use, and to prevent trademarks and patents from expiring on the timelines enshrined by law. Can we all have a moment of silence for poor Bob Iger? Maybe we can start a GoFundMe to help him out?
- hahajk 3y ago> And a whole universe of potential trademark infringements with this single two-word prompt: animated toys If you flood the market and dominate children's culture with toys from your TV shows, you absolutely cannot complain when your toys are considered iconic enough to be the generic "animated toy". These images don't replace or substitute the things they are depicting.
- pxoe 3y agothere's an easy fix. the easiest. just don't use data that you don't have the rights to use. apparently that's just impossible. "but what if we want to scrape the entire web and something makes it in anyway? see, that is impossible". well that's just saying "fuck it" and using bad data anyway. that's not an actual effort to "not use data you can't use" - there was just no way there'd be a 'rights cleared' way to use the entire web anyway. that is impossible. using a clean dataset is not impossible. it's very possible.
- oglop 3y agoSo what? I feel like I’m taking crazy pills when I read these things. You all do realize the same thing happens in your mind with those same prompts right? That’s kinda how it works. Who is surprised by this? Yeah no shit it can kinda reproduce the text it was trained on, so do I! That’s how that works. And the NYT knew for a long ass time this thing was ingesting. Literally saw this in the marketing when I signed up last year. I wasn’t shocked when I noticed I could query it about ANY math textbook I owned and it could talk with me about it. I did t bitch and gripe, I enjoyed it and have conversations. Anyway, I’m in the minority I guess. I love that I can talk with it about books and news.
- AC_8675309 3y agoSo the models overfit the training data, essentially memorizing, instead of generalizing?
- ur-whale 3y agoIt's not for generative AI that thing are about to get a lot worse. It is in fact the very notion of Copyright is breathing its last breath, and it is fantastic to be alive to see it happen.
- ultrablack 3y agoWe are all trained on copyrighted input. That is not a problem. What is a problem is if you reproduce it and try to claim copyright for that. If someone wants to create their own image of Mario in an AI, so what?
- gumballindie 3y agoWe are not machines. The argument that procedural text and image generators are similar to us is ridiculous. The issue is not whether people can generate images. The issue is ai companies stealing content and reselling it. That needs to stop.
- rvz 3y ago> The argument that procedural text and image generators are similar to us is ridiculous. Agreed. The amount of endless whataboutisms AI proponents have to continuously invent around comparing humans and AI machines as having 'similar' characteristics to justify mass copyright violation is just absolutely laughable. > The issue is ai companies stealing content and reselling it. The key point here is the 'reselling' part, without credit, attribution or permission to do so and then claiming the creation as one's own. The fact that these AI companies won't disclose their training data, tells us that they know they are in deep trouble. The so-called 'fair use' excuses isn't going to work this time. Given that Apple paid news orgs to train on their licensed data, the lawsuit with the NYT should not be a surprise for OpenAI and Microsoft (as they knew that they needed to pay for a license to access and train on the data) and will eventually end with a licensing deal with the NYT.
- noitpmeder 3y agoIt's like the ones arguing in favor of blatant AI theft have a monetary incentive to see that they succeed. Or, gasp, they are LLMs themselves.
- Log_out_ 3y agoThat sound, as if layers and layers of renteering aristocracy were forced to work again against their will.
- legendofbrando 3y agoSurely one answer is to train (or aggressively fine-tune) a new model that doesn’t (or refuses) to produce these outputs and then - as exists already, augment that model’s understanding of copyrighted material by having it Bing/Google search as a RAG process that requires the end user to log into accounts at the New York Times (and other accounts) with their paid sub. This broadly replicates the process a person could do today when they read the internet and summarize it while paying rights holders. Expensive to do but hardly the end of Generative AI or OpenAI should that be the difference between having a business or being sued out of existence. Never underestimate people who have a clear economic interest especially when their own existence is at stake.
- 1shooner 3y agoImagine a future where copyright registration involves contributing your IP to a public adversarial model, which is then a regulated layer in future generative model licensing.
- DigitallyFidget 3y agoPer United States law, imagery/art/music/text/photography generated by non-human means (such as machinery, animals, or generative AI) cannot hold copyright. https://copyright.gov/comp3/chap300/ch300-copyrightable-authorship.pdf https://copyright.gov/comp3/chap300/ch300-copyrightable-auth... Section 306 on page 7. I'm not sure how it'll hold up in law to claim copyright violations against something that wasn't created by a person. It'll really depend on the lawyers and judge's interpretation of written law. But I'm curious to see what comes of this.
- zanfr 3y agohmm then it meants generative music, as in say brian eno's experiments aren't copyrighted?
- iwontberude 3y agoI guess so! Good point.
- jimbobimbo 3y agoDid he ever use AI to generate music? As opposed to crafting and using an algorithm, in which case the computer is just an instrument, like synthesizer is.
- sgt101 3y agoSo on your interpretation if I photocopy a book and then sell the photocopies to my friends there is no infringment? I don't think so, but hey, a photocopier is a machine and it generated the book so should be ok!
- DigitallyFidget 3y agoThat isn't my interpretation, nor did I ever make that statement. That IS infact a valid definition of copyright infringement. The source material is copyrighted and you're making a literal copy of it via photocopier. I don't know how you twisted the logic on that to conclude that would not be infringement. However, it does also depend what you do with the photocopies, merely photocopying a book and keeping it privately is on par with copying a music CD as a backup. The infringement occurs when you're reusing it as your own, such as selling, publishing, or broadcasting the copyrighted material. What I stated is that generated art such as images/music/photos that are by a non human cannot be copyrighted. A photocopier isn't generating anything, it's a copy, it's replication and it isn't generating a new thing. My personal opinion is that AI generated artwork should be treated as equal to fanart when generating copyright influenced material.
- zanfr 3y agono matter how you look at it; the cat is out of the bag. OpenAI could be censored but you can't censor the opensource
- _giorgio_ 3y agoThis guy built a career around nonsensical and catastrophic endings. Everything that he sees has mysterious flaws that never happen.
- deleted 3y ago[deleted]
- aimor 3y agoI did an interesting thing and looked at how well the Llama2 models could compress text. For example, I took the first chapter of the first Harry Potter book and recorded the index of the 'correct' predicted token. The original text, compressed with 7zip (LZMA?) to about 14kB. The Llama2 encoded indexes compressed to less than 1kB. Then, of course, I can send that 1kB file around and decode the original text. (Unless the model behaves differently on different hardware, which it probably does) What I get from this is that Llama2 70B contains 93% of Harry Potter Chapter 1 within it. It's not 100% (which would mean no need to share the encoded indices) but it's still pretty significant. I want to repeat this with the entire text of some books, the example I picked isn't representative because the text is available online on the official website.
- proaralyst 3y agoWhile I don't disagree that these models seem to contain the ability to recreate copyrighted text, I don't think your conclusion holds. How well does zstd compress Harry Potter with a dictionary based on English prose? I think you'll get some impressive ratios, and I also think there's nothing infringing in this case.
- sebzim4500 3y agoCouldn't you use the same argument to reach the absurd conclusion that the 7zip source code contains the vast majority of Harry Potter? A decent control would be to compare it to similar prose that you know for a fact is not in the training data (e.g. because it was written afterwards).
- aimor 3y agoI think the same argument would have to compare 7zip's compression to some other compression algorithm. Then we can say things like "7zip is a better/worse model of human writing". And that's probably a better way to talk about this as well. You're right that a better baseline could be made using books not in the training set, to understand how much is the model learning prose and how much is learning a specific book.
- tayo42 3y ago
- ctoth 3y agoEverybody just buying into the corporate narrative that anyone can actually own these sorts of things. Who truly owns the tales of Snow White and Cinderella? These stories didn't originate with Disney; they are part of a rich tapestry of folklore passed down through generations. Disney's success was partly built on adapting these existing narratives, which were once shared and reshaped by communities over centuries. This conversation shouldn't just be about the technicalities of AI or the legalities of copyright; it should be about understanding the deep roots of our shared culture. At its core, culture is a communal property, evolving and growing through collective storytelling and reinterpretation. The current debate around AI and copyright infringement seems to overlook this fundamental aspect of cultural evolution. The algorithms might be new, but the practice of reimagining and repurposing stories is as old as humanity itself. By focusing solely on the legal implications and ignoring the historical context of cultural storytelling, we risk overlooking the essence of what it means to be a creative society. As a large human model, (no really I could probably lose some weight) I think it's just silly how we're all sort of glossing over the fact that Disney built their house of mouse on existing culture, on existing stories, and now the idea that we might actually limit the tools of cultural expression to comply with some weird outdated copyright thing is just...bonkers.
- iainctduncan 3y agoCopyright has never been based on a moral stance. It has always been determined by the lobbying power of various groups. The idea that we should dispense with it to let generative AI companies make even more money seems totally bizarre.
- logicchains 3y ago>The idea that we should dispense with it to let generative AI companies make even more money seems totally bizarre. How's that bizarre, if as you state copyright has always been based on "money makes right" not some moral stance?
- RecycledEle 3y ago> The idea that we should dispense with it [copyright] to let generative AI companies make even more money seems totally bizarre. The idea is that we should remove abuses of copyright to allow our society to move forward, and thereby continue to exist. Imagine if there was a law at the beginning of the Industrial Revolution that said when non-human labor was used, the Animal Welfare Office had veto power. Then imagine that the Animal Welfare Office declared steam engines to be immoral, and so steam engines were never used in industry, at least not in the Wester World. The Orient would eventually rise as the world's only industrial power. In the same way, if we let the copyright industry veto generative AI, it will destroy the Western World. Our students are already at a huge disadvantage compared to Chinese students who get every book ever translated into Chinese for free (except a few immoral works that they would not want to see anyway.) Those who pose an existential threat to our civilization are rent seekers who abuse copyright in the US to go beyond protecting "science and the useful arts," who seek infinite copyright terms, who grab every creative work We The People create and register lying paperwork to ensure they can steal our creative genius to enrich their cabal. If this was only a for-profit scheme, it would not be so bad. Do you remember when they Hollyweird degenerates sued a Christian company that wanted to put our G-rated versions of the movies aimed at children? The Christian company never suggested they not pay for the movies. No matter what the Christian company was willing to pay, they were not allowed to publish child-friendly versions of the movies. This proves Hollyweird's goal is to push degeneracy. The battle against abuses of copyright is a fight for Western Civilization. The fight against abuses of copyright if a fight for our souls.
- octacat 3y agoI am expecting politicians would do some nice mental gymnastics regarding regulating this. All major IT companies are doing genai now and nobody wanna hurt the companies.
- throwuwu 3y agoCopyright is fucked. Even if Open AI somehow loses this and has to delete GPT4 and their training data, the generative AI cat is so far out of the bag that it’s gone on to live a full life and have many grandkittens. It’s already easy to install and run generative models and it’s just going to get easier and the models will keep getting better. These lawsuits are futile and won’t matter in 2 years or less.
- yieldcrv 3y agoa lot worse for cloud providers hosting generative AI the models can be fine
- iainctduncan 3y agoI am constantly suprised by the amount of apologizing for generative AI infringement here. The fact that it's already being done and is a technical breakthrough is irrelevant to existing copyright law. "We are big and innovative" may hold weight with legislators, but it won't with the courts. Remember when everyone and their dog discovered sampling in the late 80's and they all thought they could get away with it because it didn't seem like infringement to the samplers? The courts had no qualms about slapping record labels for putting out records with unlicensed samples in them. Albums even got pulled off shelves while licenses were sorted out. These companies are charging for a service that returns copyrighted content, full stop. You can't do that whether you are AI or someone drawing Mario and selling the pictures on iStock, or putting out records that sample someone else's work without permission. It took a while in the case of sampling, but it sure as hell happened.
- deputy 3y ago[deleted]
- qgin 3y agoThings are about to get a lot worse for generative AI in the United States They are about to be infinitely better for generative AI in China.
- noitpmeder 3y agoChina has massive IP theft and Chile labour issues that arguably give them competitive advantages too. Should we let those slide as well?
- qgin 3y agoEach issue is unique. We can evaluate training of AI models on its own without needing to accept child labor.
- smrtinsert 3y agoThe NYTimes case is a clear one because they are delivering nearly the same content as an end product to users. The others seem like dead ends. The infringer would be the prompter, not the AI which operates more like a search engine. This is Napster all over again, what a phenomenal waste of time and money, where the artist will definitely come out with 0 at the end of it and a few corporations control everything - not to mention, there's nothing stopping anyone from releasing a tool that will crawl all spongebobs, generate your model for you and allow you to produce locally copyright infringing material it to your hearts content locally. You could drown yourself in local spongebobs.
- airstrike 3y agoI have no patriotic skin in the game, being neither American, nor European, nor Chinese, but this copyright issue seems overblown to me and like the perfect way to hand the leadership in generative AI over to China
- startupsfail 3y agoWould you prefer to live under Chino-Russia dominating the technology sector or EU-US?
- airstrike 3y agoEU-US, but were I consulted by the Chino-Russia camp, I'd say encouraging this debate is in our best interest and we should do our best to promote the issue as a real "danger"
- dmbche 3y agoHey so the problem isn't the output of the LLMs but the input - the data they are trained on is stolen (big suprise, you can't claim fair use when using something commercially, like training your LLM). The output is irrelevant. Edit1: If you want to verify this, check out all the lawsuits against AI companies : it's always about using their copywritten goods. Any discussion about the output is to talk about the amount of damage done to the copyright holder, not if damage exists or not.
- kromem 3y agoHere's one of the senior legal peeps at the EFF who has litigated IP cases talking about the issue: https://www.eff.org/deeplinks/2023/04/how-we-think-about-copyright-and-ai-art-0 https://www.eff.org/deeplinks/2023/04/how-we-think-about-cop... It's not as clear cut as you think it is.
- wslh 3y agoWhile different, I find this discussion about AI and copyrights as an evolution of the war that never was: Google/FB converting in the portal/proxy for content and while it is not generative AI you can find copyrighted images just using Google Images or as an snippet in the normal search engine. I mention Google because it is the de facto monopoly but this applies to a lot of aggregators. I know we are talking about different technologies but it seems all these people were very silent and find some opportunity in having this war with OpenAI (not an endorsement) but not fighting others. I am not making an statement about the morals of AI and aggregators/search engines (super interesting discussion that in a way was happening for long) but I am surprised that organizations are "just" waking up. It seems they just see it is a much simple and cheap fight.
- theamk 3y agoThe thing with Google is it is super trivial to exclude your text - tag on page, header on server, etc.. So all the conversations about google "stealing" context always seemed pretty silly to me. Compared to that AI offers no way to opt out, which is a big difference.
- dmbche 3y agoPersonal use of copywritten material is fine - there is no breach of copyright when you download a picture from Google for yourself. If you use it commercially then there is breach. Uploading copywritten content is a breach of copyright as well, even without commercial use. Google/Facebook are hosting and giving access to a bunch of media, which might or might not be copywritten - it's the individuals problem.They make.money from ads, not from the content. AI companies stole copywritten media to train their commercial LLM, sell them or their products and make profit. I don't think it's the same.
- gfodor 3y agoGary Marcus is the master of AI FUD
- goertzen 3y agoNo they are not. This is a negotiation tactic by the NYT to drive up the licensing price. Period. The Napster/Music Industry analogy has no resemblance to this situation. The only meaningful question that might be answered as a result of this is, what permission and access rights do crawlers have to content that is publicly and legally available.
- 8organicbits 3y agoSurely there's a meaningful question about copying and distributing content verbatim, which GPT has been shown to do.
- sgt101 3y agoAlso the use of the content as per provision on the web. NYT is paywalled - you have to agree to a license to access it, there are exclusions in that agreement that I don't understand but I think may be important in this discussion!
- CuriouslyC 3y agoNot really. Models are a device capable of producing protected content given some input contortions. So are Xerox machines.
- 8organicbits 3y agoIf I Xerox'd a book and sold copies to people I'm clearly violating copyright. I'm not sure I follow.
- CuriouslyC 3y agoNobody has given Xerox an injunction against researching or building copiers because you can copy books and sell them.
- 8organicbits 3y ago
- SKILNER 3y agoI don't understand the glee so many people have over this. I love being able to use Generative AI tools. How is it different than if I asked a person to draw these pictures for me? I know someone will gleefully clobber this question with a legal answer, but God, let's move forward, hunh?
- wharvle 3y agoA bunch of rich people are raiding a little bit of work, each, from a whole bunch of people, then walling it off so they can get richer. I’d not have a problem with this, personally, if their models were as available as the stuff they took from others. Instead it’s take, take, take… now wait a minute, that pile of loot I stole is mine!
- dang 3y agoRelated ongoing thread: NY times is asking that all LLMs trained on Times data be destroyed - https://news.ycombinator.com/item?id=38816944 https://news.ycombinator.com/item?id=38816944 - Dec 2023 (93 comments) Also: NY Times copyright suit wants OpenAI to delete all GPT instances - https://news.ycombinator.com/item?id=38790255 https://news.ycombinator.com/item?id=38790255 - Dec 2023 (870 comments) NYT sues OpenAI, Microsoft over 'millions of articles' used to train ChatGPT - https://news.ycombinator.com/item?id=38784194 https://news.ycombinator.com/item?id=38784194 - Dec 2023 (84 comments) The New York Times is suing OpenAI and Microsoft for copyright infringement - https://news.ycombinator.com/item?id=38781941 https://news.ycombinator.com/item?id=38781941 - Dec 2023 (861 comments) The Times Sues OpenAI and Microsoft Over A.I.’s Use of Copyrighted Work - https://news.ycombinator.com/item?id=38781863 https://news.ycombinator.com/item?id=38781863 - Dec 2023 (11 comments)
- amai 3y agoShould the NYT not sue https://commoncrawl.org/ https://commoncrawl.org/ ? OpenAI just used the data from commoncrawl for training.
- noitpmeder 3y agoIs that true? Has OpenAI revealed exactly what is in their training set?
- karmakaze 3y agoIt shouldn't matter how the images/etc are created. The problem comes about when it's used as an original work by the person that's doing so. Imagine instead of AI/ML, we have a mechanical-turk-like service that produces output from descriptions. The service makes no claims that the generated outputs are not similar to any copyrighted works. The only claim the service makes is that they themselves claim no copyright on the output. It's then up to the user of the service to determine if the output is suitable for their intended use. Whether such a service itself is legal is a separate matter. For that matter, say you outsourced the artwork to a person who again gave you infringing work. The user of that output is still in violation. With AI/ML we're basically outsourcing to a 'service' that is known to sometimes output copyrighted work so with the user knowing that, are responsible for fair usage.
- wayeq 3y agoWe need to figure out how to ever so gradually move toward a post-copyright economy.
- karmakaze 3y agoThe real 'problem' is how do we navigate the present and near future where much more than physical labor is being automated? This is where we need sustainable solutions. The rough road on the way should also be smoothed out so as not to disrupt so many lives, but it's good to keep a perspective what and why we're doing these things.
- jlnthws 3y agoWe could get inspiration from the case of the record industry against Napster, or cabs VS Uber. Both parties are somehow abusing their position, but the world is moving on. Rent seeking is probably not an absolute source of wealth after all.
- RecycledEle 3y agoRent seeking should be a capital crime.
- RecycledEle 3y agoIf we get rid of unconstitutional copyrights in the US, this ges away. Recall that according to the US Constitution, copyright can only be on on "science and the useful arts." Alternately, we could restore a reasonable limit to the duration of copyrights, like 14 years.
- shkkmo 3y agoIt seems like this article makes a basic copyright mistake. I don't see any evidence that these are " reproductions" of source material like since no source image is linked to compare. Instead, these are derivative works. We already have a flourishing culter of derivitave works, such as fan art that exist in various shades of legal greyness. Some derivative works are fair use, some are not. The position of the Author here seems to be that generative AI should not be capable of creating any derivitave works, or should only be able to do so it it can accurately identify which are fair use and which aren't (which seems like an impossibly tall bar.) This stance seem like a giant attack on fair use that significantly expands the power of copyright. To me, the takeaway from this is different. This makes clear that there is currently a risk when using AI generated art that you could end up unintentionally creating and publishing a derivative work unintentionally and thus without evaluating if that work constitues fair use.
- tim333 3y agoThey are just going to have to inform the AI in some sense of the current copyright situation and ask it not to infringe. It's the same for human writers. If you are writing an article for Wikipedia say, you should read relevant source articles and then rewrite in a way that isn't a copy and paste beyond a few words.
- noitpmeder 3y agoOk I'll bite. Let's assume you've informed the current models about copyright and asked them not to infringe... What happens when they continue to do so.
- ofslidingfeet 3y agoI'm still waiting for people to figure out the whole point of an automated process is that it behaves the same way each time.
- KETpXDDzR 3y agoI'd expect "Open"AI et al to lobby heavily towards an "AI-generated content is excluded from copyright infringement". I think it's possible that they'll introduce a "generative AI" tax. Charge x cents per generated text/image and distribute the fund to all media companies. In Germany you pay some amount extra on top of the sales price of anything that can store data (CX, DVD, USB sticks, HDDs, ...). This is then distributed to all companies that could be impacted by software piracy. I'm still not sure if that's legal considering the Geneva convention disallows collective punishment.
- zer0c00ler 3y ago[dead]
- appplication 3y agoThere are an alarming number of responses seemingly completely unaware of the core thrust of the article (and NYT lawsuit). ChatGPT was able to reproduce and publish significant portions of NYT articles, completely verbatim for hundred-to-thousand word stretches. It’s not derivative work. We’re way past that. NYT has an exceptionally strong case here and anyone arguing about the merits of copyright is way off the mark. This court case is not going single-handedly to undo copyright. OpenAI has very little going for them other than “this is new, how were we to know it could do this”. So knowing that, the currently trained models are in a very sticky situation. Further, I don’t see NYT settling. The implications are too large, and if they settle with OpenAI, they will have a similar case pop up with every other model. And every other publisher of digital content with have a similarly merited case. This is an inflection point for generative AI, and it’s looking like it will be either much more expensive or much more limited than we originally thought. A side effect of this: I am predicting that we will start to see a rise in “pirate” models. Models who eschew all legality, who are trained in a distributed fashion, and whose weights are published not by corporations but by collectives (e.g. torrent models). There is a good chance we see these surpass the official “well behaved” models in effectiveness. It will be an interesting next few years to see this play out.
- RestlessAPI 3y agoSuch a thing happened with DALLE, Midjourney, and Stable Diffusion. Stable Diffusion, when used to its fullest with thing like Control Net and LoRAs, blows the pants off of other proprietary models.
- benlivengood 3y agoMy guess is that OpenAI will be able to basically copy Google/YouYube on this and offer a system like content-ID. Specifically, ChatGPT doesn't reproduce copyrighted works by default; only by request/action of a third party user much like YouTube serving whatever videos people upload. It wasn't the intent of OpenAI to infringe copyright and in fact a lot of or most researchers believed the models were not overfitted enough to reproduce significant portions of arbitrary works.
- NemoNobody 3y ago
- 8note 3y ago"from classic sci-fi movie" How could you put that as the prompt without intending to infringe? Anything pulled from a classic sci-fi movie would be infringement. The term droid is also star wars specific? Id consider the "red soda" one as grounds that the Coca-Cola brand has become generic and that it's synonymous with soda. Same thing with Mario too. There is so much non-nintendo content made featuring Mario the plumber that you could get that without training directly on Nintendo's artwork
- sjducb 3y agoI think it’s a question of what counts as publication. I think that an AI model is analogous to an employee. Imagine I ask my employee to write an article, and they just copy an existing one from the times. That’s plagiarism and bad work, not copyright infringement. If I then decide to publish the plagiarised article, then I have committed copyright infringement. I once ran into this exact problem with a human. I hired a designer to make some artwork for an app. When I launched the app it turned out that the human had just copied the artwork from another game. It’s my problem that I hired an idiot, and my problem that my app was infringing the copyright of another app. (We redesigned the graphics very quickly)
- null_point 3y agoI suspect this may delay some short term progress by creating pressure on AI labs to train their models from data curated or synthesized in a way that is contentious of copyright law. There is already troves of data that are fair game for training, but even "corrupted" data sets can probably be used if used intelligently. We've already seen examples of new models effectively being trained off of GPT-4. That approach with filters for copyrighted material might allow for data that is sufficiently "scrambled". Not to say building such a filter is definitely easy, but seems plausible.