10 ms·
'Impossible' to create AI tools like ChatGPT without copyrighted material, OpenA
- dsq 3y agoWell, you cant raise a human without 'copyrighted' (or at least already existing) material in the form of parental norms, ethical codes, schoolbooks, pre existing humans like parents, grandparents, etc. We dont spring like Athena from the forehead of Zeus. Heinlein, in his book "The Moon is a Harsh Mistress" made the interesting point that an AI needs love in order to become a full entity. Love needs to come from from somewhere.
- belter 3y agoIn other words...Your rights to the content you created, are disturbing the road to my IPO?
- leereeves 3y agoIf we judge ChatGPT by the same standards we judge humans, it would still be in a lot of trouble for copyright violations and plagarism. (Not that I think we should do so - ChatGPT is not AGI.)
- edgyquant 3y agoPeople really need to stop trying to make this argument. Even if you were right about this, ChatGPT IS NOT A HUMAN! It literally can’t be any clearer than that, society and its laws exist by and for humans and human rights and informational norms do not apply to mechanical statistical models. And no, I don’t care if you believe humans are “totally just a similar type of model.”
- logicchains 3y agoPeople are not saying ChatGPT is a human, they're saying it learns like a human so its learning should be legally treated like a human learning. Laws are based on philosophy, and none of the philosophies upon which western notions of human rights are based rely on the fact that it's a biological human as justification, rather they rely on arguments that apply equally to any conscious self-aware being. Saying that only humans should ever have rights because they're biological makes as much logical sense as only Caucasians should ever have rights because their skin is white.
- RandomLensman 3y agoWe generally don't treat machines as human, why make an exception here? Animals are also not treated as humans, so the bar for a machine to clear is very very high to get that treatment, which then would include criminal matters, too, I suppose.
- logicchains 3y agoMost machines don't learn like humans; these machines do. It'll be even more apparent when we put LLMs in mechanical bodies and allow them to update weights in real time, which is only a few years away given there aren't any real technological impediments. It's stupid to suggest that it should be illegal for an LLM in a body to watch a copyrighted movie because it could memorise and quote it verbatim, when that's what humans do.
- RandomLensman 3y agoSo what if that were true - no reason to treat them as human because they are not human (or even close to human) presently.
- logicchains 3y agoWhat reason is there to treat them differently just because one's made of silicon and the other of carbon? Sounds like just plain ol' racism.
- RandomLensman 3y agoDo we treat a car as human for the purposes of locomotion? Should we treat certain fluorocarbons as human because they can work a bit like blood? There is nothing racist in treating a machine like a machine. Just because a machine replicates or approximates something a human does, doesn't mean it is a human or should be treated as one.
- 3y ago
- EMIRELADERO 3y agoSure, but I don't see how that's a good response to the argument. It's not about "the machine's right to do something", but "my right to make the machine do something". That's what's at play, and what can be allowed or disallowed.
- RandomLensman 3y agoFair point. Your right to do something can clearly differ from your right to make a machine do something. A simple example would be you being allowed to walk down footpath, but letting your car traverse that path isn't allowed. If we are talking hypothetical new rules, not immediately obvious to me what is optimal here, including differentiation of various use cases.
- close04 3y agoMore importantly the machine is usually allowed to do less than the human. Otherwise it would become an easier avenue for abuse. Whenever you can't do something just build a machine to do it, and the machine's capabilities are ever expanding. Maybe the starting point for legislating "machines" should be to take human capability as a baseline, make some humans eventually accountable, and go from there. If your machine is learning truly like a human (takes decades to learn then output results, forgets things, only one in a thousand can output something useful at the end of the learning process, etc.) and the "owners" take full accountability for what it does then why not like a human? But this reminds me of an old joke. A man comes into an inn and asks how much for a thimble of water. The inn keeper says a thimble is free. So he proceeds to ask for 1000 thimbles.
- mjbeswick 3y agoThat's your opinion, but right now he law doesn't cover training AI models. Do you think that content created people is truly original?
- edgyquant 3y agoWhy does this matter at all? We’re talking about statistical models not humans
- chrisjj 3y agoSure it does. It binds the human trainers.
- anotherhue 3y agoIf the models had a total lifetime output similar to a human we wouldn't be worrying about this. Instead they're rather fast, and that makes them something different.
- skepticATX 3y agoThe materials that you are referencing were all created for the benefit of humans. The creators want humans to learn from them.
- strogonoff 3y agoIf you cannot create an LLM without copyrighted material, then your duty as an SV-based tech company sitting on billions of USD is to hire an army of lawyers and license that material. Go on, do things that don’t scale.
- throwaway29812 3y ago> Go on, do things that don’t scale. Nice.
- addandsubtract 3y agoRight? I could create the next Netflix that has ALL the content, if it wasn't for those pesky copyright laws.
- pierat 3y agoYou can do it right now, in Jellyfin. And with Sonarr, Radarr, Lidarr, Jellyseerr, Prowlarr, flaresolverr, VPN, and qbittorrent - you can automate a request->watch pipeline for your users. And if you use distributed block storage, you can make this scale pretty damn big. (Prowlarr scrapes torrent sites for viable torrents. Uses Flaresolverr to bypass cloudflare shit. Sonarr Lidarr Radarr - categorizes and prepares downloads for movies, tv shows, and music and submits them to qbittorrent. Jellyseerr is a request dashboard that works alongside Jellyfin that allows users to request stuff. VPN/Qbittorrent is a combined tool thats used as a download-endpoint for the whole thing.)
- w3ll_w3ll_w3ll 3y agoIs this made up?
- pierat 3y agoWhat do you mean "made up"? The Arrs suite is definitely a thing, as are the other tools I talk about. There's a reference docker that automates this whole toolchain https://github.com/AdrienPoupa/docker-compose-nas https://github.com/AdrienPoupa/docker-compose-nas Whats nice is that Prowlarr will scrape all the torrent sites for goodies to index, and your users in jellyfin can just request, and it just shows up. You can even have a discord channel and the tools will announce when stuff's done. Piracy is 10000x better than paying for continually worsening streaming (pile of shit). Edit: And frankly, the billionaires copy with little ramifications, and add paywalls. I'll copy as well. The big difference is I do it for free.
- breadwinner 3y agoAI training is no different than a human reading copyrighted material and training his biological brain. In neither case is the "brain" allowed to regurgitate source material in its original form. And as long as that doesn't happen there is no copyright violation.
- deleted 3y ago[deleted]
- happytoexplain 3y agoYou're missing the critical part where that human is also a genius alien super-prodigy who can recreate or recombine any IPs in seconds and then we made him available on the internet to everybody.
- anotherhue 3y agoExactly, if I can memorise books and I sell access to call me and have me recite books then that's infringement.
- breadwinner 3y agoWhat if you're not "reciting" books, but using the knowledge gained from books to teach large lecture-halls full of college students? Is that a copyright violation, even a slight violation?
- otikik 3y agoBut they do in fact produce copies of their training materials. This has been demonstrated both with image generation (a targeted query will reproduce a very similar image to the one used in training) and in source code generation (a machine produced the source code of a program that it had been trained over).
- simonw 3y agoThe NYT lawsuit includes 100 examples of GPT-4 being able to produce copies of its training data, not using RAG: https://nytco-assets.nytimes.com/2023/12/Lawsuit-Document-dkt-1-68-Ex-J.pdf https://nytco-assets.nytimes.com/2023/12/Lawsuit-Document-dk...
- WhackyIdeas 3y agoIt’s also impossible to pirate movies without pirating movies. It’s one rule for them and another for the rest of us.
- deleted 3y ago[deleted]
- speak_plainly 3y agoRelevant precedent from Japan: https://analyticsindiamag.com/japan-sets-the-precedent-for-ai-copyright/ https://analyticsindiamag.com/japan-sets-the-precedent-for-a...
- xyzal 3y agoThe legal entity owning the brain should be responsible for that brains output. That is, you should be responsible for the output of your brain and OpenAI for the output of its LLMs. Seems simple to me.
- logicchains 3y agoWhy shouldn't the entity using the brain be responsible, since they're the one directing it?
- chrisjj 3y agoTell me where the LLM user directed the LLM to slurp (C) works.
- logicchains 3y agoThe NYT prompted it with sections of copyright articles to get it to reproduce the rest of the copyrighted article. Which is pretty clearly deliberately trying to get it to reproduce a copyright article.
- Snow_Falls 3y ago... Which shows that its capable of regurgitating copyrighted material. What does it matter what the method is? People can deliberately use the thing for copyright infringement. Besides, proving that it can be done on purpose is just the easiest method, your really think it won't happen by accident?
- dcow 3y agoIf you prompted me with: "<blank> for a Klondike." and I responded with "What would you do for a Klondike." then I've recreated copyrighted material. Does it make sense to sue me? I've learned the phrase organically, and I may have never seen the original content, only learned it 2nd hand via word of mouth. If I charged you for my time does it matter? Note I fully agree that OpenAI should acquire the appropriate licenses to use all of the content it uses to train its models. However I'm not as clear on whether anybody should be able to place limits on or modulate/attenuate the output once the input has been appropriately licensed and responsibly consumed.
- seydor 3y ago..but also it is very possible to limit what people do with ChatGPT output material
- mattgreenrocks 3y agoTech company in front of investors: "WE ARE THE ALPHA OF OUR INDUSTRY AND NOTHING WILL STAND IN OUR WAY" Tech company in front of government/media: "Our job is so hard, please let us be exempt from laws because...progress!"
- 6gvONxR4sf7o 3y agoIt’s funny to see arguments about how this brand new product at the frontier of science (that society did fine without) is only possible this one particular way. It’s silly. We just learned to make it at all, we don’t know what is and isn’t possible. In 30 years (no time at all, in terms of the legal implications of whatever laws are set up for this stuff), we’ll have a completely different conception of chatbots.
- mensetmanusman 3y agoThis is true, which is why it is destiny for next gen AI companies to not emerge from western countries.
- mikewarot 3y agoClearly our copyright terms are too long, and as a result the public domain is insufficient to train an AI, which seems to suggest it is also insufficient for a number of other uses, including the public good.