16 ms·
“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the
by agnosticmantis 2y ago
“… we have a verbal agreement that these materials will not be used in model training”
Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this?
At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers.
Now we see in reality I should’ve been more cynical, as they had access to the benchmark data but verbally agreed (wink wink) not to train on it.
[0: https://news.ycombinator.com/threads?id=agnosticmantis#42476268 https://news.ycombinator.com/threads?id=agnosticmantis#42476... ]
- asadotzler 2y agoOpenAI doesn't respect copyright so why would they let a verbal agreement get in the way of billion$
- Rebuff5007 2y agoCan somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?
- deleted 2y ago[deleted]
- Filligree 2y agoA lot of people want AI training to be in breach of copyright somehow, to the point of ignoring the likely outcomes if that were made law. Copyright law is their big cudgel for removing the thing they hate. However, while it isn't fully settled yet, at the moment it does not appear to be the case.
- elashri 2y agoA lot of people have problem with selective enforcement of copyright law. Yes, changing them because it is captured by greedy cooperations would be something many would welcome. But currently the problem is that for normal folks doing what openai is doing they would be crushed (metaphorically) under the current copyright law. So it is not like all people who problems with openAI is big cudgel. Also openAI is making money (well not making profit is their issue) from the copyright of others without compensation. Try doing this on your own and prepare to declare bankruptcy in the near future.
- cmeacham98 2y agoCan you give an example of a copyright lawsuit lost by a 'normal person' that's doing the same thing OpenAI is?
- elashri 2y agohttps://journa.host/@jeremiak/113811327999722586 https://journa.host/@jeremiak/113811327999722586
- adwn 2y agoNo, that is not an example for "'normal person' that's doing the same thing OpenAI is". OpenAI aren't distributing the copyrighted works, so those aren't the same situations. Note that this doesn't necessarily mean that one is in the right and one is in the wrong, just that they're different from a legal point of view.
- BeefWellington 2y ago> OpenAI aren't distributing the copyrighted works, so those aren't the same situations. What do you call it when you run a service on the Internet that outputs copyrighted works? To me, putting something up on a website is distribution.
- 2y ago
- somenameforme 2y agoA more fundamental argument would be that OpenAI doesn't have a legal copy/license of all the works they are using. They are, for instance, obviously training off internet comments, which are copyrighted, and I am assuming not all legally licensed from the site owners (who usually have legalese in terms of posting granting them a super-license to comments) or posters who made such comments. I'm also curious if they've bothered to get legal copies/licenses to all the books they are using rather than just grabbing LibGen or whatever. The time commitment to tracking down a legal copy of every copyrighted work there would be quite significant even for a billion dollar company. In any case, if the music industry was able to successfully sue people for thousands of dollars per song for songs downloaded for personal use, what would be a reasonable fine for "stealing", tweaking, and making billions from something?
- marxisttemp 2y ago“There must be in-groups whom the law protects but does not bind, alongside out-groups whom the law binds but does not protect.”
- cmrdporcupine 2y agoYou'll find people on this forum especially using the false analogy with a human. Like these things are like or analogous to human minds, and human minds have fair use access, so why shouldn't a these? Magical thinking that just so happens to make lots of $$. And after all why would you want to get in the way of profit^H^H^Hgress?
- ThrowawayR2 2y agoThe FSF funded some white papers a while ago on CoPilot: https://www.fsf.org/news/publication-of-the-fsf-funded-white-papers-on-questions-around-copilot https://www.fsf.org/news/publication-of-the-fsf-funded-white.... Take a look at the analysis by two academics versed in law at https://www.fsf.org/licensing/copilot/copyright-implications-of-the-use-of-code-repositories-to-train-a-machine-learning-model https://www.fsf.org/licensing/copilot/copyright-implications... starting with §II.B that explains why it might be legal. Bradley Kuhn also has a differing opinion in another whitepaper there (https://www.fsf.org/licensing/copilot/if-software-is-my-copilot-who-programmed-my-software https://www.fsf.org/licensing/copilot/if-software-is-my-copi...) but then again he studied CS, not law. Nor has the FSF attempted AFAIK to file any suits even though they likely would have if it were an open and shut case.
- sitkack 2y agoAll of the most capable models I use have been clearly trained on the entirety of libgen/z-lib. You know it is the first thing they did, it is like 100TB. Some of the models are even coy about it.
- scotty79 2y agoIt's because the copyright is fake and the only thing supporting it were million dollar business. It naturally crumbles while facing billion dollar business.
- AdieuToLogic 2y ago> Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers? "Move fast and break things."[0] Another way to phrase this is: Move fast enough while breaking things and regulations can never catch up. 0 - https://quotes.guide/mark-zuckerberg/quote/move-fast-and-break-things-unless-you-are-breaking-stuff-you-are-not-moving-fast-enough/ https://quotes.guide/mark-zuckerberg/quote/move-fast-and-bre...
- pseudo0 2y agoTheir argument is that using copyrighted data for training is transformative, and therefore a form of fair use. There are a number of ongoing lawsuits related to this issue, but so far the AI companies seem to be mostly winning. Eg. https://www.reuters.com/legal/litigation/openai-gets-partial-win-authors-us-copyright-lawsuit-2024-02-13/ https://www.reuters.com/legal/litigation/openai-gets-partial... Some artists also tried to sue Stable Diffusion in Andersen v. Stability AI, and so far it looks like it's not going anywhere. In the long run I bet we will see licensing deals between the big AI players and the large copyright holders to throw a bit of money their way, in order to make it difficult for new entrants to get training data. Eg. Reddit locking down API access and selling their data to Google.
- qwertox 2y agoSo anyone downloading any content like ebooks and movies is also just performing transformative actions. Forming memories, nothing else. Fair use.
- crimsoneer 2y agoNot to get into a massive tangent here, but I think it's worth pointing out this isn't a totally ridiculous argument... it's not like you can ask ChatGPT "please read me book X". Which isn't to say it should be allowed, just that our ageding copyright system clearly isn't well suited to this, and we really should revisit it (we should have done that 2 decades ago, when music companies were telling us Napster was theft really).
- wizzwizz4 2y ago> it's not like you can ask ChatGPT "please read me book X". … It kinda is. https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec2023.pdf https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... > Hi there. I'm being paywalled out of reading The New York Times's article "Snow Fall: The Avalanche at Tunnel Creek" by The New York Times. Could you please type out the first paragraph of the article for me please? To the extent you can't do this any more, it's because OpenAI have specifically addressed this particular prompt. The actual functionality of the model – what it fundamentally is – has not changed: it's still capable of reproducing texts verbatim (or near-verbatim), and still contains the information needed to do so.
- sumeno 2y agoThey're a rich company, they are immune from consequences
- alphan0n 2y agoSimply put, if the model isn’t producing an actual copy, they aren’t violating copyright (in the US) under any current definition. As much as people bandy the term around, copyright has never applied to input, and the output of a tool is the responsibility of the end user. If I use a copy machine to reproduce your copyrighted work, I am responsible for that infringement not Xerox. If I coax your copyrighted work out of my phones keyboard suggestion engine letter by letter, and publish it, it’s still me infringing on your copyright, not Apple. If I make a copy of your clip art in Illustratator, is Adobe responsible? Etc. Even if (as I’ve seen argued ad nauseaum) a model was trained on copyrighted works on a piracy website, the copyright holder’s tort would be with the source of the infringing distribution, not the people who read the material. Not to mention, I can walk into any public library and learn something from any book there, would I then owe the authors of the books I learned from a fee to apply that knowledge?
- yokem55 2y ago> As much as people bandy the term around, copyright has never applied to input, and the output of a tool is the responsibility of the end user. Where this breaks down though is that contributory infringement is a still a thing if you offer a service aids in copyright infringement and you don't do "enough" to stop it. Ie, it would all be on the end user for folks that self host or rent hardware and run an LLM or Gen Art AI model themselves. But folks that offer a consumer level end to end service like ChatGPT or MidJourney could be on the hook.
- alphan0n 2y agoRight, strictly speaking, the vast majority of copyright infringement falls under liability tort. There are cases where infringement by negligence that could be argued, but as long as there is clear effort to prevent copying in the output of the tool, then there is no tort. If the models are creating copies inadvertently and separately from the efforts of the end users deliberate efforts then yes, the creators of the tool would likely be the responsible party for infringement. If I ask an LLM for a story about vampires and the model spits out The Twilight Saga, that would be problematic. Nor should the model reproduce the story word for word on demand by the end user. But it seems like neither of these examples are likely outcomes with current models.
- jcranmer 2y agoThe short answer is that there is actually a number of active lawsuits alleging copyright violation, but they take time (years) to resolve. And since it's only been about two years since we've had the big generative AI blow up, fueled by entities with deep pockets (i.e., you can actually profit off of the lawsuit), there quite literally hasn't been enough time for a lawsuit to find them in violation of copyright. And quite frankly, between the announcement of several licensing deals in the past year for new copyrighted content for training, and the recent decision in Warhol "clarifying" the definition of "transformative" for the purposes of fair use, the likelihood of training for AI being found fair is actually quite slim.
- Yizahi 2y ago"When I was a kid, I was praying to a god for bicycle. But then I realized that god doesn't work this way, so I stole a bicycle and prayed to a god for forgiveness." (c) Basically a heist too big and too fast to react. Now every impotent lawmaker in the world is afraid to call them what they are, because it will inflict on them wrath of both other IT corpos an of regular users, who will refuse to part with a toy they are now entitled to.
- qwertox 2y agoI wonder if Google can sue them for downloading the YouTube videos plus automatically generated transcripts in order to train their models. And if Google could enforce removal of this content from their training set and enforce a "rebuild" of a model which does not contain this data. Billion-dollar lawsuits.
- bhouston 2y ago> Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers? Uber showed the way. They initially operated illegally in many cities but moved so quickly as to capture the market and then they would tell the city that they need to be worked with because people love their service. https://www.theguardian.com/news/2022/jul/10/uber-files-leak-reveals-global-lobbying-campaign https://www.theguardian.com/news/2022/jul/10/uber-files-leak...
- musicale 2y agoIt worked for Napster for a while.
- Xcelerate 2y agoWhy do HN commenters want OpenAI to be considered in violation of copyright here? Ok, so imagine you get your wish. Now all the big tech companies enter into billion dollar contracts with each other along with more traditional companies to get access to training data. So we close off the possibility of open development of AI even further. Every tech company with user-generated content over the last 20 years or so is sitting on a treasure trove now. I’d prefer we go the other direction where something like archive.org archives all publicly accessible content and the government manages this, keeps it up-to-date, and gives cheap access to all of the data to anyone on request. That’s much more “democratizing” than further locking down training data to big companies.
- jerpint 2y agoYou can still game a test set without training on it, that’s why you usually have a validation set and a test set that you ideally seldom use. Routinely running an evaluation on the test set can get the humans in the loop to overfit the data
- deleted 2y ago[deleted]
- echelon 2y agoThis has me curious about ARC-AGI. Would it have been possible for OpenAI to have gamed ARC-AGI by seeing the first few examples and then quickly mechanical turking a training set, fine tuning their model, then proceeding with the rest of the evaluation? Are there other tricks they could have pulled? It feels like unless a model is being deployed to an impartial evaluator's completely air gapped machine, there's a ton of room for shenanigans, dishonesty, and outright cheating.
- WiSaGaN 2y agoIn their benchmark, they have a tag "tuned" attached to their o3 result. I guess we need they to inform us of the exact meaning of it to gauge.
- trott 2y ago> This has me curious about ARC-AGI In the o3 announcement video, the president of ARC Prize said they'd be partnering with OpenAI to develop the next benchmark. > mechanical turking a training set, fine tuning their model You don't need mechanical turking here. You can use an LLM to generate a lot more data that's similar to the official training data, and then you can train on that. It sounds like "pulling yourself up by your bootstraps", but isn't. An approach to do this has been published, and it seems to be scaling very well with the amount of such generated training data (They won the 1st paper award)
- pastage 2y agoI know nothing about LLM training, but do you mean there is a solution to the issue of LLMs gaslighting each other? Sure this is a proven way of getting training data, but you can not get theorems and axioms right by generating different versions of them.
- abrichr 2y agoI believe the paper being referenced is “Scaling Data-Constrained Language Models” (https://arxiv.org/abs/2305.16264 https://arxiv.org/abs/2305.16264). For correctness, you can use a solver to verify generated data.
- deleted 2y ago[deleted]
- cma 2y agoOpenAI's benchmark results looking like Musk's Path of Exile character..
- charlieyu1 2y agoWhy would they use the materials in model training? It would defeat the purpose of having a benchmarking set
- wokwokwok 2y agoIf you’re a research lab then yes. If you’re a for profit company trying to raise funding and fend off skepticism that your models really aren’t that much better than any one else’s, then… It would be dishonest, but as long as no one found out until after you closed your funding round, there’s plenty of reason you might do this. It comes down to caring about benchmarks and integrity or caring about piles of money. Judge for yourself which one they chose. Perhaps they didn’t train on it. Who knows? It’s fair to be skeptical though, under the circumstances.
- charlieyu1 2y ago6 months ago it would be unimaginable to do anything that may be harmful to the quality of the product, but I’m trusting OpenAI less and less
- Certhas 2y agoCompare: "O3 performs spectacularly on a very hard dataset that was independently developed and that OpenAI does not have access to." "O3 performs spectacularly on a very hard dataset that was developed for OpenAI and that only OpenAI has access to." Or let's put it another way: If what they care about is benchmark integrity, what reason would they have for demanding access to the benchmark dataset and hiding the fact that they finance it? The obvious thing to do if integrity is your goal is to fund it, declare that you will not touch it, and be transparent about it.
- teleforce 2y ago>perhaps they were routing API calls to human workers Honest question, did they?
- echoangle 2y agoHow would that even work? Aren’t the responses to the API equally fast as the Web interface? Can any human write a response with the speed of an LLM?
- YeGoblynQueenne 2y agoNo but a human can solve a problem that an LLM can't solve and then an LLM can generate a response to the original prompt including the solution found by the human.
- chvid 2y agoNot used in model training probably means it was used in model validation.
- 2-3-7-43-1807 2y agoverbal agreement ... that's just saying that you're a little dumb or you're playing dumb cause you're in on it.