8 ms·
What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost
by carlosdp 3y ago
What you described is entirely fair use, actually.
Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)
- hn_throwaway_99 3y ago> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use. 2. People should familiarize themselves with the four factors of fair use determination. In particular, if a work is purely derivative of a source work and substantially negatively impacts the market for the original work, it's very likely to not be considered fair use. A great overview is https://fairuse.stanford.edu/overview/fair-use/four-factors/ https://fairuse.stanford.edu/overview/fair-use/four-factors/
- NegativeK 3y ago> suddenly everyone's a copyright lawyer Roll back 20+ years ago on Slashdot and you'll see the exact same thing. Copyright has been a hot button issue on the internet for decades. People end up thinking (rightly or wrongly) that they understand it without being a lawyer.
- Teever 3y agoOne of my biggest gripes is a somewhat adjacent issue where everyone thinks they're an American copyright lawyer and that American copyright law is universal. It's very possible that the example provided above is an example of fair use in some country, and that the website offering that service could be hosted there.
- whoknowsidont 3y ago> that they understand it without being a lawyer. Quite literally, not even the lawyers or courts understand it. This is very much a "learn as you go" exercise for humanity in general at this point in time.
- Mattbrown7531 3y agoIt seems like everything in tech is in the learn as you go phase. Everything is changing so rapidly that there can’t be experts. Just people that are able to adapt quickly. I only see this phenomenon speeding up. Strange times.
- Vicinity9635 3y agoLegality aside I think copyright of digital things in the digital age is a net negative to humanity.
- matheusmoreira 3y agoCompletely agree. Copyright should be abolished. All intellectual work is information, information is just bits and bits are just numbers. It's quite simply delusional to believe you can own numbers in the 21st century, the age of information and ubiquitous globally networked pocket supercomputers. This is just a felony contempt of business model issue. Computers invalidated their business models and they're doing everything they possibly can to hang on for dear life. Society needs to move on already.
- paulryanrogers 3y agoThis goes too far. Digital media are not only long series of numbers. They are often difficult-to-create expressions in image, video, and even interactive forms; regardless of their serialization format. Books are just strings of letters, yet copyright has still been useful to increase the volume and utility of books. All that said, I do find the life+70y an absurdly long time.
- blehn 3y agoThat seems like a great way to destroy what is left of art as we know it. Anna Karenina is just numbers. In The Mood For Love is just numbers. Right. What do you propose is the business model for artists in the absence of copyright?
- matheusmoreira 3y agoI propose getting paid before doing the work for the actual labor of creation. Crowdfunding, patronage, comissions, sponsorships all seem like ethical ways to get things done sustainably. That way creators get paid before they work, not after. We must strengthen these business models that don't depend on artificial scarcity because this number selling nonsense was over the second computers were invented. It's as dumb as asserting that you need permission to use memcpy or the mov CPU instruction.
- bomewish 3y agoThis seems a bit disingenuous. Lawyers DISAGREE on this stuff (as we will see in this case) and a court will decide the reality by fiat.
- deleted 3y ago[deleted]
- paulddraper 3y ago> if a work is purely derivative of a source work CliffNotes, Wikipedia, etc. have huge quantities of summarized copyrighted work.
- caesil 3y agoSummarization generally isn't considered a derivative work. https://en.wikipedia.org/wiki/Derivative_work https://en.wikipedia.org/wiki/Derivative_work
- btilly 3y agoFirst, you missed the "and". Do CliffNotes, Wikipedia, etc. substantially impact the market for the original work? For example CliffNotes does not - people who buy the CliffNotes version typically already have the original work as well (for example from coursework). And Wikipedia may well do more to interest people in the original work than to replace it. Second, you ignored the "purely derivative" bit. You have to look at to what extent the use is derivative or transformative. See https://en.wikipedia.org/wiki/Transformative_use https://en.wikipedia.org/wiki/Transformative_use for a bit about that. (Note, this is a legal term defined by various precedents. OpenAI can't just argue, "Turning it into an LLM is a transform, so it is transformative!") Since CliffNotes is educational and Wikipedia is nonprofit, it is relatively easy for both to qualify as transformative. As a result your response underscores the point that was made. There are a lot of shades of grey. You really can't just seize on a couple of phrases and key points, then jump straight to the answer. You have to understand how the courts will decide, and then accept that there is an actual judgment call whose outcome depends on the judge judging. (I'm not a lawyer, but I have had excessive exposure to them in the past.)
- joegahona 3y ago> people who buy the CliffNotes version typically already have the original work as well (for example from coursework) Is there data that supports this? I’d be interested to know what % of people who buy a Cliffs Notes have already _bought_ the original.
- 3y ago
- shkkmo 3y ago> Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use. You're the one presenting unfounded claims with confidence here. There is well established case law about not being able to copyright facts. If you are actually fully paraphrasing a presentation of facts / ideas and not just altering a couple of words here and there, then there is a very strong case for non-infringement.
- hn_throwaway_99 3y ago> You're the one presenting unfounded claims with confidence here. No, I'm not. On the contrary, I'm really looking forward to this case because I believe it will be a great test of a bunch of concepts that are totally novel in the world of copyright law as it applies to generative AI. The only things I am presenting with confidence are: 1. That anyone who declares that something is unambiguously fair use (or, contrarily, unambiguously infringing) is likely wrong. There is simply too much latitude by judges, and there have certainly been cases where a ruling went one way, only to be overturned on appeal. 2. While I certainly have an opinion on how I think this case will be decided, I'm not presenting that with unwarranted confidence. Instead, I linked that great article on the 4 factors of fair use determination because it's clear to me lots of people are saying "fair use!" on one side or the other with no understanding of the factors judges must actually consider when making a determination.
- deleted 3y ago[deleted]
- shkkmo 3y agoYou seem to be shifting the topic of this thread. The GP comment is about paraphrasing news articles while I don't see anything in the NYT lawsuit about paraphrasing. Rather, the NYT is concerned with exact reproduction or near exact reproduction. I too am very curious about the outcome of this case and wouldn't care bet either way on the outcome. I do have an opinion on what precedent would be better for our society but that doesn't mean I think that outcome is more likely. However, none of that matters in this particular thread. There are well established precedents about paraphrasing news articles and they do not support the claim you made
- caesil 3y ago>if a work is purely derivative of a source work This is the weakest part of the case(s) against OpenAI. "Derivative work" is a legal term of art meaning a direct adaptation, like writing a screenplay of a book or translating a book into another language. NYT has a stronger case than Sarah Silverman here because they can show actual 'memorized' text rather than just summarization, but given that those memorizations are a) an unintended failure mode of the training process, and b) from an older version of the model that has been updated to no longer regurgitate memorized text, it's not really clear how in current form GPT could possibly be considered a derivative work.
- avidiax 3y agoA question is whether the new model still intrinsically embeds the source text, but this is later filtered in the output, or if it no longer embeds the text at all. The latter is more defensible.
- skygazer 3y agoI would think an existing model could bootstrap a copyright free training corpus by completely rewriting/paraphrasing copyrighted material with semantic fidelity for training of the next model to completely eliminate memorization of copyrighted works. That might pose an interesting obstacle to copyright challenges, bootstrapping your way into a clean room. Although, tweaking the architecture to either eliminate memorization, or eliminate high fidelity reproduction of verbatim training data seems far more expedient and less costly.
- dchichkov 3y ago"Transformative" seems to fit a lot more that "Derivative". On the other hand, it's understandable why NYT is worried. OpenAI itself says that occupations like: Writers and Authors, Web and Digital Interface Designers, News Analysts, Reporters, and Journalists, Proofreaders and Copy Markers are "90-100% exposed" to what OpenAI is building.
- paledot 3y ago
- bsenftner 3y agoI personally appreciate the semi truck sized loophole that is satire. One can include an entire copy written work within one's own work as long as the treatment of that other copy written work is parody / satire. This is a provision of US copyright law put in place to protect political satire, which can be anything, because politics is everything.
- Forgeties79 3y agoI would say it is arguable that is fair use, but the whole thing about fair use is that it is a defense, not a type of license or something you can preemptively apply. So whether or not it will be protected under fair use is actually not determined yet. In fact I would say that’s the entire debate here, right? I have worked on many documentaries and any time we said “fair use” internally what we were implicitly saying is “nobody will come after us because they know that we are probably safe under fair use if this escalated.“ But again, we could never preemptively apply it. We were just anticipating potential conflict and gauging how likely it was to occur.
- h1fra 3y agoit's fair use if you don't make money from your project no?
- JohnFen 3y agoNo. In the US, whether or not you make money has little to do with whether or not your use qualifies as "fair use".
- semiquaver 3y agoWhy do you say that? Commercial vs noncommercial use is a primary factor in the “purpose” prong of the fair use balancing test and a significant one in the “market effects” prong. That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.
- freejazz 3y agoBecause anyone that is familiar with fair use knows that the purpose prong and the commerciality aspect of it is not one of the more important prongs of the fair use analysis, whereas transformation is. Transformation adjusts what is a purpose that falls under fair use. Did you read Warhol??
- semiquaver 3y agoYes. Warhol is an example where the commercial nature of the secondary use was the deciding factor in its failure to pass the purpose prong. > In sum, if an original work and secondary use share the same or highly similar purposes, and the secondary use is commercial, the first fair use factor is likely to weigh against fair use, absent some other justification for copying. (P4). It’s very likely that a noncommercial secondary use would have passed under the reasoning in Warhol. I don’t understand the point you’re trying to make.
- freejazz 3y ago
- freejazz 3y ago>What you described is entirely fair use, actually. Based upon what? You think other publishers use NYTimes articles for free without license?
- ummonk 3y agoHe's talking about citing and quoting NYTimes articles, not republishing them verbatim. That said, it's very different if you're a publication that sometimes cites reporting from other publications vs. a website exclusively dedicated to indexing and summarizing NYTimes articles.
- fennecbutt 3y agoI couldn't get gpt to quote an actual nyt article no matter how hard I tried...it just hallucinated in the general style of a news article. Presumably, if it can remember at least a paragraph or two of each article, then surely the same would be true of any text it ingested and the model size would approach the dataset size (probably actually much larger). I don't believe this is the case at all, even searching around, I've not found any good recent examples of it regurgitating copyrighted text verbatim. It's cool to hate AI stuff if you're a creative atm. But gotta love those generative/algorithm based PS brushes, that's still real art! "Indeed, the opening paragraph of "A Game of Thrones" by George R.R. Martin, with the chapter titled "Bran," starts as follows: "The morning had dawned clear and cold, with a crispness that hinted" And then it cuts off, whether that's because OAI now have an oh shit filter or just the model had access to the first page or publicly available articles quoting the first line, I'm not sure. I tried other chapters and random sections and it could get a sentence or two right but then hallucinated; what's more likely NYT and GRRM? That your works are being reproduced verbatim? Or that Facebook, YouTube descriptions, fan tumblrs and hell, the publicly available and multiple GoT related wikis that include a variety of passages from the books were used as training data?
- ummonk 3y agoI don't think it's necessarily true that model size would need to be larger than dataset size. It's theoretically possible that the model encoding achieves significantly better compression than DEFLATE or GZIP or whatever compression algorithm is used to store the dataset itself.
- Powdering7082 3y agoDo you have some examples & are you sure they don't pay licensing fees to NYT?
- _the_inflator 3y agoA lawyer starts a conclusion like this: "It could be fair use if conditions a, b, and c are met. Condition a means..." ;)
- object-a 3y agoSourcing, quoting, and linking is covered in the NYTimes content policy under fair use. See: https://help.nytimes.com/hc/en-us/articles/115014891408-Obtaining-and-using-Times-content https://help.nytimes.com/hc/en-us/articles/115014891408-Obta... I think what wouldn't be covered is reproducing substantial portions of an article, especially if it's done without attribution. Tier 2 publications that fully reprint NYT or AP/Reuters articles are usually doing this via a paid News Service or Content License. See: https://nytlicensing.com/content/new-york-times-news-service/ https://nytlicensing.com/content/new-york-times-news-service...
- laszlojamf 3y agoIsn't that the thing though? If they credit the source, it's fine, but ChatGPT usually doesn't
- windexh8er 3y agoCorrect. But those 2nd tier sources don't have the NYT copy verbatim. Do you really think the US NFL, as an example, would let OpenAI use all of its recorded games as a way to train some new GenAI game framework to build better American Football games? No. All that material is copyright. Public media is going to move to a very awkward era of ownership and licensing because all of these large companies looking to make a buck off public data sets are doing very little to make the economic model less one sided. I hope the NYT prevails here, personally. Models will (and are) currently tainted by data they should not contain and for longer term privacy concerns this needs to be addressed early and have significant consequences or we're headed towards a world where this type of technology will make our ad-targeted world seem like a much more manageable past.