14 ms·
>> In the proposal, OpenAI also said the U.S. needs “a copyright strategy that promotes the freedom to learn” and on “preserving American AI models’ ability to
by dchichkov 2y ago
>> In the proposal, OpenAI also said the U.S. needs “a copyright strategy that promotes the freedom to learn” and on “preserving American AI models’ ability to learn from copyrighted material.”
Perhaps also symmetric "freedom to learn" from OpenAI models, with some provisions / naming convention? U.S. labs are limited in this way, while labs in China are not.
- thrance 2y agoThey meant "freedom to learn [through backpropagation]" probably. Companies like this were allowed to siphon the free work of billions of people over centuries and they still want more.
- sega_sai 2y agoI like how this "freedom to learn" should apply to models, but not real people..
- IncreasePosts 2y agoAre there certain books that federal law prevents you from reading? Which ones? Maybe terrorist manuals and some child pornography, but what else?
- TheSoftwareGuy 2y agoIt already applies to real people, doesn't it? I.e. if you read a book, you're not allowed to start printing and selling copies of that book without permission of the copyright owner, but if you learn something from that book you can use that knowledge, just like a model could.
- m1el 2y agowhen it comes to real people, they get sued into oblivion for downloading copyrighted content, even for the purpose of learning. but when facebook & openai do it, at a much larger scale, suddenly the laws must be changed.
- ryoshu 2y agoCase in point - https://en.wikipedia.org/wiki/Aaron_Swartz https://en.wikipedia.org/wiki/Aaron_Swartz
- JumpCrisscross 2y agoSwartz wasn’t “downloading copyrighted content…for the purpose of learning,” he was downloading with the intent to distribute. That doesn’t justify how he was treated. But it’s not analogous to the limited argument for LLMs that don’t regurgitate the copyrighted content.
- Terretta 2y ago> when it comes to real people, they get sued into oblivion for downloading copyrighted content, even for the purpose of learning. Really? Or do they get sued for sharing as in republishing without transformation? Arguably a URL providing copyrighted content, is you offering a xerox machine. It seems most "sued into oblivion" are the reshare problem, not the get one for myself problem.
- conjectures 2y agoIt does apply to people? When you read a copy of a book, you can't be sued for making a copy of the book in the synapses of your brain. Now, if you have eidetic memory and write out large chunks of the book from memory and publish them, that's what you could be sued for.
- triceratops 2y ago> When you read a copy of a book They're not talking about reading a book FFS. You absolutely can be sued for illegally obtaining a copy of the book.
- deleted 2y ago[deleted]
- tsimionescu 2y ago
- triceratops 2y agoCan I download a book without paying for it, and print copies of it? Stash copies in my bathroom, the gym, my office, my bedroom etc. to basically have a copy on hand to study from whenever I have some free time? What about movies and music?
- Terretta 2y agoIs the book online and accessible to your eyeballs through your open standards client tool, such that you can learn from seeing it?
- triceratops 2y agoMost books aren't. Unless you pay for them.
- codedokode 2y agoLet's say Windows is downloadable from Microsoft website. Can you use it for free in your company to save on buying a license? Is it ok to use illegal copies of works in a business?
- ajross 2y ago> Can I download a book without paying for it, and print copies of it? No, but you can read a book, learn its contents, and then write and publish your own book to teach the information to others. The operation of an AI is rather closer to that than it is to copyright violation. "Should" there be protections against AI training? Maybe! But copyright law as it stands is woefully inadequate to the task, and IMHO a lot of people aren't really treating with this. We need a functioning government to write well-considered laws for the benefit of all here. We'll see what we get.
- triceratops 2y agoBut I can't legally obtain the book to read and learn from without me (or a library) paying for it. Let's start there first.
- echelon 2y agoIf models can learn for free, then the models (training code, inference code, training data, weights) should also be free. No copyright for anybody. And if you sell the outputs of your model that you trained on free content, you shouldn't be able to hide behind trade secret.
- crorella 2y ago> just like a model could It is not remotely the same, the companies training the models are stealing the content from the internet and then profiting from it when they charge for the use of those models.
- Terretta 2y ago> the companies training the models are stealing the content from the internet Are you stealing a billboard when you see and remember it? The notion that consuming the web is "stealing" needs to stop.
- crorella 2y agoWe are not taking about billboards here, we are talking about copyrighted works, like books. If you want to do mental gymnastics and call "consuming" the web the act of downloading books without paying for them, then go ahead, but don't pretend the rest will buy your delusion.
- Terretta 2y agoOn the contrary, even telling people which billboards are posted about what, and how to get to them to look at them, is "how it works". But the courts will get to clarify (in today's news): https://www.reuters.com/legal/news-corp-sued-by-brave-software-google-search-engine-rival-2025-03-13/ https://www.reuters.com/legal/news-corp-sued-by-brave-softwa...
- llamaimperative 2y agoThe question is whether it destroys the incentive to produce the work. That is the entire point of copyright and patent law. LLMs do indeed significantly reduce the incentive to produce original work.
- codedokode 2y agoAre you stealing when using a pirated software to run a billion-dollar business?
- 2y ago
- simion314 2y ago>you can use that knowledge, Did OpenAI bought one copy of each book, or did they legaly borowed athe books and documents ? if you copy paste rom books and claim is your content you are plagiarizing. LLMs were provent to copy paste trained content so now what? Should only big Tech be excluded from plagiarizing ?
- pier25 2y ago> just like a model could Not really. You can't multiply yourself a million times to produce content at an industrial scale.
- alabastervlog 2y agoThis is why I think my array of hard drives full of movies isn't piracy. My server just learned about those movies and can tell me about them, is all. Just like a person!
- tsimionescu 2y agoIt doesn't, a real person can't legally obtain a copy of a copyrighted work without paying the copyright holder for it. This is what OpenAI is asking for: they don't want to pay for a single copy of a single book, and still they want to train their models on every single book in history (and song, and movie, and painting, and code base, and anything else they can get their hands on).
- bee_rider 2y agoThese AI models are just obviously new things. They aren’t people, so any analogy about learning from the training material and selling your new skills is off base. On the other hand, they aren’t just a copy of the training content, and whether the process that creates the weights is sufficiently transformative as to create a new work is… what’s up for debate, right? Anyway I wish people would stop making these analogies. There isn’t a law covering AI models yet. It is a big industry at this point, and the lack of clarity seems like something we’d expect everybody (legislators and industry) to want to rectify.
- amelius 2y agoTotally agree. Except the current administration probably will interpret things the way they see fit ...
- codedokode 2y agoModel cannot "learn" because it is not a human. What happens is a human obtains "a free copy" of a copyrighted work, processes it using a machine and sells the result.
- bee_rider 2y ago> Model cannot "learn" because it is not a human. Sure, that’s why don’t like the analogy. > What happens is a human obtains "a free copy" of a copyrighted work, processes it using a machine and sells the result. Right, so for example it is pretty common to snip up small bits of songs and to use in other songs (sampling). Maybe that could be an example of somewhere to start? But, these ML models seem quite different, I guess because the “samples” are much smaller and usually not individually identifiable. And really the model encodes information about trends in the sources… I dunno. I still think we need a new law.
- aiono 2y agoCan I pirate books to train myself?
- amelius 2y agoDo you know Numerical Recipes in C? This discussion reminds me of it.
- sidewndr46 2y agoAnd when I "learn" a verbatim copy of pages of that book, then write those pages out in Microsoft Word & sell those pages its legal?
- DebtDeflation 2y agoEven moreso, it only applies to initial model training by companies like OpenAI not other companies using those models to generate synthetic data to train their own models.
- binarymax 2y agoYeah it’s crazy. I also suspect they might not be confident in their defense from the NYT lawsuit - if they’re found in fault then it’s going to be trouble.
- logsr 2y agoIt is hard to see how a court could decide that copyright does not apply to training LLMs without completely collapsing the entire legal structure for intellectual property. Conceptually, AI basically zeros out existing IP, and makes the AI the only IP that has any value. It is hard to imagine large rights holders and courts accepting that. The likely outcome is that courts rule against LLM creators/providers and they eventually have to settle on licensing fees with large corporate copyright holders similar to YouTube. Unlike YouTube though, this would open up LLM companies to class action lawsuits from the general public, and so it could be a much worse outcome for them.
- diego_sandoval 2y agoI would assume that the request is for it to apply to models in the way that it currently applies to humans. If a human buys a movie, he can watch it and learn about its contents, and then talk about those contents, and he can create a similar movie with a similar theme. If OpenAI buys a movie and shows it to their model, it's unclear whether the model can talk about the contents of the movie and create a similar movie with a similar theme.
- m1el 2y agosomehow, I suspect openai didn't "buy" all of the articles, books, websites they crawled and torrented.
- mitthrowaway2 2y agoIs OpenAI buying the movie, or just taking it? Since "buying" a movie (as it currently applies to humans) is just buying a limited license to it for private viewing, can't the copyright holder opt to limit the $4.99 license terms to human viewing, and charge $4999 for an AI training license? Or OpenAI could buy movies the way Disney does, by buying the actual copyright to the film.
- htrp 2y ago>Since "buying" a movie (as it currently applies to humans) is just buying a limited license to it for private viewing, can't the copyright holder opt to limit the $4.99 license terms to human viewing, and charge $4999 for an AI training license? the Reddit data licensing model
- da_chicken 2y ago> Since "buying" a movie is just buying a license to it, can't the copyright holder opt to limit the $4.99 license terms to human viewing, and charge $4999 for an AI training license? That's exactly what already happens currently. Buying a movie on DVD doesn't give you the right to present it for hundreds of people. You need to pay for a public performance license or commercial licence. This is why a TV network or movie theatre can't just buy a DVD at Walmart and then show the movie as often as it likes. Copyright doesn't just grant exclusive distribution rights. It grants exclusive use rights as well, and permits the owner to control how their work is used. Since AI rights are not granted by any existing licenses, and license terms generally reserve any rights not explicitly specified, feeding copyrighted works into an AI data model is a reserved right of the owner.
- voytec 2y agoThis is basically "allow us to steal others' IP". It's hard not to treat Altman like a common thief.
- kranke155 2y agoNot only that The model gets to use training data of all humans. But if you use the model as training data OAI will say you’re infringing T&Cs
- deleted 2y ago[deleted]
- taurath 2y agoIt still warps my brain, they’ve taken trillions of dollars of industry and made a product worth billions by stealing it. IP is practically the basis of the economy, and these models warp and obfuscate ownership of everything, like a giant reset button on who can hold knowledge. It wouldn’t be legal, or allowed if tech wasn't seen as the growth path of our economy. It’s a hell of a needle to thread and it’s unlikely that anyone will ever again be able to model from data so open.
- woah 2y ago"IP" is a very new concept in our culture and completely absent in other cultures. It was invented to prevent verbatim reprints of books, but even so, the publishing industry existed for hundreds of years before then. It's been expanded greatly in the past 50 years. Acting like copyright is some natural law of the universe that LLMs are upending simply because they can learn from written texts is silly. If you want to argue that it should be radically expanded to the point that not only a work, but even the ideas and knowledge contained in that work should be censored and restricted, fine. But at least have the honesty to admit that this is a radical new expansion for a body of law that has already been radically expanded relatively recently.
- mullingitover 2y ago> It was invented to prevent verbatim reprints of books It was also invented to keep the publishing houses under control and keep them from papering the land in anti-crown propaganda (like the stuff that fueled the civil war in England and got Charles I beheaded). Probably one of the biggest brewing fights will be whether the models are free to tell the truth or whether they'll be mouthpieces for the ruling class. As long as they play ball with the powers that be, I predict copyrights won't be a problem at all for the chosen winners.
- jsemrau 2y ago"mouthpieces for the ruling class" That's actually a great point. Judging from the current state of media, there is a clear momentum to take sides in moral arguments. Maybe the standard for models need to be a fair use clause?
- deleted 2y ago[deleted]
- EGreg 2y agoGearing up for a fight between the two major industries based on exploitative business models: Copyright cartels (RIAA, MPAA) that monetized young artists without paying them much at all [1], vs the AI megalomaniacs who took all the work for free and used Kenyans at $2 an hour [2] so that they can raise "$7 trillion" for their AI infrastructure [1] https://www.reddit.com/r/LetsTalkMusic/comments/1fzyr0u/artistsbands_destroyed_by_the_music_industry_how/ https://www.reddit.com/r/LetsTalkMusic/comments/1fzyr0u/arti... [2] https://time.com/6247678/openai-chatgpt-kenya-workers/ https://time.com/6247678/openai-chatgpt-kenya-workers/
- Bjorkbat 2y agoCan't believe I'm actually rooting for the copyright cartels in this fight. But that does make me think, that in a sane society with a functional legislature I wouldn't have to pick a dog in this fight. I'd have have enough faith in lawmakers and the political process to pursue a path towards copyright reform that reigns in abuses from both AI companies and megacorp rightsholders Alas, for now I'm hoping that aforementioned megacorps sue OpenAI into a painful lesson.
- visarga 2y ago> Can't believe I'm actually rooting for the copyright cartels in this fight. The same megacorps are suing Internet Archive for their collection of 78rpm records. These guys would rather see art orphaned and die.
- disgruntledphd2 2y agoYup, we live in a pretty depressing world. More generally the best we can hope for us to discourage concentrated power, both in government and corporate forms.
- __loam 2y agoThey're suing Internet Archive because IA scanned a bunch of copyrighted books to put online for free (e: without even attempting to get permission to do so) then refused to take them down when they got a C&D lol. IA is putting the whole project at risk so they can do literal copyright infringement with no consequences.
- blitzar 2y agoI should have "freedom to learn" about any Tesla in the showroom, any F-35 I see laying around an airbase or the contents of anyone in the governments bank account.
- NoOn3 2y agoAccording to this scheme, if you find a bug and can read the bank's data, then you can use it as you want.
- JonChesterfield 2y agoNope, have to feed it into an llm first, afterwards it's legitimate.
- NoOn3 2y agoNo need for a LLM. Humans always have their own neural networks in their heads. :)
- deleted 2y ago[deleted]
- seanmcdirmid 2y agoChinese AI must implement socialist values by law, but law is a much more fluid fuzzy thing in China than in the USA (although the USA seems to be moving away from rule of law recently).
- sva_ 2y ago> Chinese AI must implement socialist values by law I don't doubt it but am interested to read a source? I know the models can't talk about things like Tiananmen Square 1989, but what does 'implementing socialist values by law' look like?
- doctorwho42 2y agoSocialism and freedom of speech aren't mutually exclusive
- seanmcdirmid 2y agohttps://www.cnbc.com/2024/07/18/chinese-regulators-begin-testing-genai-models-on-socialist-values.html https://www.cnbc.com/2024/07/18/chinese-regulators-begin-tes... "Socialist values" is literally the language that China used in announcing this. Here is a recent article from a Chinese source: https://www.globaltimes.cn/page/202503/1329537.shtml https://www.globaltimes.cn/page/202503/1329537.shtml Although censorship isn't mentioned specifically, it is definitely 99% of what they are focused on (the other 1% being scams). China practices Rule by law, not Rule of law, so you know...they'll know its bad when they see it, so model providers will exercise extreme self censorship (which is already true for social network providers).
- janalsncm 2y ago> China practices Rule by law, not Rule of law In practice the US is less different than you imply. For the vast majority of Americans, being sued is a punishment in and of itself due to the prohibitive costs of hiring a lawyer. In the US we have a right to a “speedy” trial but there are many people sitting in jail now because they can’t afford the bail get out. Speedy could mean months. I say this because when we constantly fall so far short of our ideals, one begins to question if those are really our ideals.
- samstave 2y ago[flagged]
- cadamsdotcom 2y agoCan you expand your post and explain why?
- eunos 2y agoMust be the rumours that DeepSeek has million worths of GPU think and their claim of relatively cheap training is a psyop
- janalsncm 2y agoThe pod was good apart from starting/spreading the rumor that high numbers of “bill to Singapore” was evidence that China was circumventing GPU import bans.
- samstave 2y agoDont look at it as such, mayhaps; Look it at literally who will have GPU dominance in future. (obv who will hit Qbit at scale... but we are at this scale now - and control currently is controlled by policy, then bits, then Qbits.) Remember, we are witnessing the "Wouldnt it be cool if..?" CyberPunk manifestations of our Cyberpunk Readings of youth? ((I buildt a bunch of shit that spied on you because I read NeuroMancer, and thought wouldnt it be cool if..." And then I helped build Ono Sendai throughout my career...
- cscurmudgeon 2y ago[dead]
- 999900000999 2y agoCan this extend to every kid sued by the record industry for downloading a few songs. Have we all been transported to bizzaro land? Different rules for billion dollar corps I guess.
- somenameforme 2y agoThose cases did very poorly whenever they actually went to court (well at least also including the ones that were summarily dismissed by the courts, meaning they didn't technically make it to court). They were much more of a mafia style shakedown than an actual legal enforcement effort. Same rules, but people are a lot less inclined to defend themselves because the cost of loss was seen as too high to even risk it.