9 ms·
Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM
by dissident_coder 3y ago
Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States.
And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can do other than put everything you create behind some kind of authentication wall but even then it’s only a matter of time until it leaks anyway.
Pandora’s box is really open, we need to figure out how to live in a world with these systems because it’s an un winnable arms race where only bad actors will benefit from everyone else being neutered by regulation. Especially with the massive pace of open source innovation in this space.
We’re in a “mutually assured destruction” situation now, but instead of bombs the weapon is information.
- llm_nerd 3y agoI don't think they're looking to prevent the inevitable, but rather see a target with a fat wallet from which a lot of money can be extracted. I'm not saying this in a negative way, but much of the "this is outrageous!" reaction to AI hasn't been about the building of models, but rather the realization that a few players are arguably getting very rich on those models so other people want their piece of the action.
- dissident_coder 3y agoIf NYT wins this, then there is going to be a massive push for payouts from basically everyone ever…I don’t see that wallet being fat for long.
- noitpmeder 3y agoIf they are determined to have broken the law then they should absolute be made to pay damages to aggrieved parties (now, determining if they did and who those parties are is an entirely unknown can of worms)
- throw_nbvc1234 3y agoThe data will have to become more curated. Exclusivity deals will probably become a thing too. Good data will be worth the money and hassle; garbage (or meh) data won't.
- alexey-salmin 3y agoIf LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.
- logicchains 3y agoBy that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.
- password4321 3y ago> the copyright holder of every library book gets paid
- nullindividual 3y agoCopyright holders do get paid for library copies, in the US.
- exitb 3y agoYou make it seem as if the copyright holder is making more money on a library book, than on one sold in retail, which does not appear to be the case in the US.
- willseth 3y agoThe library pays for the books and the copyright holder gets paid. This is no different from buying a book retail, which you can read and share with family and friends after reading, or sell it, where it can be read again and sold again. The book is the product, not a license for one person to access the book.
- deleted 3y ago[deleted]
- 3y ago
- halukakin 3y agoIf this is inevitable (and I'm not saying it's not), who will produce high quality news content?
- hcurtiss 3y agoAI. And, I fear, it will be good.
- epc 3y agoCurious how AI gets the raw information if there are no reporters nor newspapers. Does AI go to meetings or interview politicians?
- hcurtiss 3y agoI can certainly imagine email correspondence. Even audio interviews. You're right that it seems at least presently AI is less likely to earn confidences. But I don't know how far off the movie "Her" actually is.
- hfhdjdks 3y agoSo Chinese LLMs are bad actors, but USA LLMs are the good guys? I don't see it that way, but I'm sure from an American perspective that how it seems.
- Salgat 3y agoWhat? This is about whether one country wants to cede a massive economic advantage to another country.
- bradchris 3y agoOn the other hand, you could also argue that if AI takes all financial incentives from professionals to produce original works, then the AI will lose out on quality material to train on and become worse. Unless your argument is there’s no need for anything else created by humanity, everything worth reading has already been written, and humanity has peaked and everyone should stop? Like all things, it’s about finding a balance. American, or any other, AI isn’t free from the global system which exists around us— capitalism.
- logicchains 3y ago>financial incentives from professionals to produce original works People produce countless volumes of unpaid works of art and fiction purely for the joy of doing so; that's not going to change in future.
- sensanaty 3y agoAnecdotal but I know lots of creatives (and by creatives I also include some devs) who've stopped publishing anything publicly because of various AI companies just stealing everything they can get their hands on. They don't mind sharing their work for free to individuals or hell, to a large group of individuals and even companies, but AIs really take it to a whole different level in their eyes. Whether this is a trend that will accelerate or even make a dent in the grand scheme of things, who knows, but at least in my circle of friends a lot of people are against AI companies (which is basically == M$) being able to get away with their shenanigans.
- gumballindie 3y agoThis argument is moot. Just because some countries - see china - steal intellectual property it doesnt mean we should. There are rules to the games we play specifically so we dont end up like them.
- ndsipa_pomu 3y agoIt's impossible to "steal" intellectual property without some kind of mind wiping device.
- noitpmeder 3y agoYou must have used that device if you're making that argument in good faith.
- ndsipa_pomu 3y agoOkay, so how is it possible to take and deprive the author of their original? The correct term would be "unauthorised copying".
- b4ke 3y agoOk, let’s address this from the standpoint of a node in the network of the thoughtscape. A denizen of the “inter”net, and also a victim of the exploitive nature of artists. Media amalgamated power by farming the lives of “common” people for content, and attempt to use that content to manage lives of both the commons and unique, under the auspice of entertainmet. Which in and of itself is obviously a narrative convention which infers implied consent (id ask to what facetiously). Keepsake of the gods if you will… We are discussing these systems as though they are new (ai and the like, not the apple of iOS), they are not… this is an obfuscation of the actual theft that’s been taking place (against us by us, not others). There is something about reaping what you sow written down somewhere, just gotta find it. -mic
- skwirl 3y agoThe word ‘moot’ does not mean what you think it means.
- tempodox 3y agoThe war on drugs has also been unwinnable from the start and yet they built an economy on top of it, with entire agencies and a prison industry. When it comes to the fabrication and exploitation of illegality, unwinnability may be a feature, not a bug.
- thebradbain 3y agoAny piece of pie deemed too big for one person to eat will be split accordingly. I don’t think NYT, or any other industry, for that matter knows AI isn’t going away: in fact, they likely prefer it doesn’t, so long as they can get a slice of that pie. That’s what the WGA and SAG struck over, and won protections ensuring AI enhanced scripts or shows will not interfere with their royalties, for example.
- ndsipa_pomu 3y agoThis suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.
- phone8675309 3y agoWhat incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?
- ndsipa_pomu 3y agoI would guess the monetisation is going to be limited to either subscriptions or advertising if your reputation allows people to especially value your curation of facts/reporting etc. The big issue with LLMs is the lack of reliability - it might be accurate or it might be an hallucination. Personally, I think it would be a lot simpler if the internet was declared a non-copyright zone for sites that aren't paywalled as there's already a legal grey area as viewing a site invariably involves copying it. Maybe we'll end up with publishers introducing traps/paper towns like mapmakers are prone to do. That way, if an LLM reproduces the false "fact", it'll be obvious where they got it from.
- asvitkine 3y agoTo have a positive impact on the world? Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with their data and everyone working there is still getting paid for their work...
- ethanbond 3y agoOh thank goodness we can rely on charity for our information economy > Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with [NYT’s] data… That’s exactly the question. They are claiming it is destroying their business, which is pretty much self-evident given all the people in here defending the convenience of OpenAI’s product: they’re getting the fruits of NYTimes’ labor without paying for it in eyeballs or dollars. That’s the entire value prop of putting this particular data into the LLMs.
- godzillabrennus 3y agoThey probably didn’t start with a lawsuit. They started asking for royalties. They probably didn’t get an offer they thought was fair and reasonable so they sued. These media businesses have shareholders and employees to protect. They need to try and survive this technological shift. The internet destroyed their profitability but AI threatens to remove their value proposition.
- deleted 3y ago[deleted]
- user432678 3y agoSorry, how exactly LLM threatens NYT? Are people supposed to generate news themselves? Or like wait a year or so before NYT articles are consumed by LMMs?
- denton-scratch 3y agoNYT doesn't just publish "news" as in what happened yesterday; they also publish analysis, reviews of books and films, history, biography and so on. That's why people cite NYT articles from decades ago.
- mc32 3y agoI’m ambivalent. On the one hand, they should realize they are one of today’s horse carriage manufacturers. They’ll only survive in very narrow realms (someone has to build the Central Park horse carriages still), but they will be miniscule in size and importance. On the other hand, LLMs should observe copyright and not be immune to copyright.
- lionkor 3y agoAn LLM in Russia can commit the same crime in Russia, and get sued in Russia. No idea about China, but I know Russia has a working legal system.
- ceejayoz 3y agoFor some definitions of “working”.
- lionkor 3y agoWorking enough that people and companies there exist, live, and are to some degree successful, yes. I've visited multiple times in the past few years and I found it to be pretty normal
- ceejayoz 3y ago“Works on my machine!” Navalny probably has a different opinion. There isn’t a country on the planet that doesn’t have people and companies. That doesn’t mean they all have functional legal systems.
- meowface 3y agoMy understanding is they have one of the most corrupt and unjust legal systems of the developed countries.
- sausagefeet 3y agoWhat's the actionable advice here? US regulation should be the lowest common denominator of all countries one considers in competition? Certainly Chinese and Russian LLMs could vacuum up all the information. China already cares little about copyright and trademark, should they stop being enforced in the US? My opinion is that the US should do things that are consistent with their laws. I don't think a Chinese or Russian LLM is much of a concern in terms of this specific aspect, because if they want to operate in the US they still need to operate legally in the US.
- woodruffw 3y agoAll of this can be true (I don’t think it necessarily is, but for the sake of argument), but it’s legally irrelevant: the court is not going to decide copyright infringement cases based on geopolitical doctrines. Courts don’t decide cases based on whether infringement can occur again, they decide them based on the individual facts of the case. Or equivalently: the fact that someone will be murdered in the future does not imply that your local DA should not try their current murder cases.
- skwirl 3y agoThe issue here is that the case law is not settled at all and there is no clear consensus on whether OpenAI is violating any copyright laws. In novel cases like this where the courts essentially have to invent new legal doctrines, I think the implications of the decision carries a tremendous amount of weight with the judges and justices who have to make that decision.
- cush 3y agoI see a complete economic collapse unless creators start getting paid both for their data upfront, and paid royalties when their data is used in an LLM response
- amanaplanacanal 3y agoCopyright doesn’t protect data, it only protects expression.
- cush 3y agoWhile I didn't say anything about copyright (obviously our current copyright laws are completely ill-equipped to handle how LLMs work), feel free to replace data with whatever you like. writing, art, music, etc. It's all the same.
- gwervc 3y agoAccess to ressources is hardly a new problem: when I was an NLP graduate student about a decade ago a teacher of us had scrapped (and continued to do so) a major newspaper for years to make a corpus. The legality of that was questionable at best, yet it was used in academic paper and a subset for training. The same is equally applicable to image: Google got rich in part by making illegal copies of whatever image he could find. Existing regulations could be updated to include ML model but that won't stop bad or big enough actors to do what they want. > We’re in a “mutually assured destruction” situation now No, we aren't. Very good spam generators aren't comparable to mass destruction weapons.
- logicchains 3y agoTrying to prevent AI from learning from copyrighted content would look completely stupid in a decade or two when we have AIs that are just as capable as humans, but solely due to being made of silicon rather than carbon are banned from reading any copyrighted material. Banning a synthetic brain from studying copyrighted content just because it could later recite some of that content is as stupid as banning a biological person from studying copyrighted content because it could later quote from it verbatim.
- 34679 3y agoWe have this now with humans. I've been in a lifelong sruggle for knowledge and tools that I can afford.
- tovej 3y agoIt's not exactly a synthetic brain though, is it? LLMs are more like lookup tables for the texts they're trained on. We will not have "AIs as capable as humans" in a couple decades. AIs will keep being tools used by humans. If you use copyrighted texts as input to a digital transformation, that's vopyright infringement. It's essentially the same situation as sampling in music, and imo the same solutions can be applied here: e.g. licenses with royalties.
- Aurornis 3y ago> Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. Foreign companies can be barred from selling infringing products in the United States. Russian and Chinese consumers are less interested in English-language articles. I can’t really get behind the argument that we need to let LLM companies use any material they want because other countries (with other languages, no less) might not have the same restrictions. If you want some examples of LLMs held back by regulations, look into some of the examinations of how Chinese LLMs are clearly trained to avoid answering certain topics that their government deems sensitive.
- logicchains 3y ago>Chinese LLMs are clearly trained to avoid answering certain topics that their government deems sensitive But they're not; you can download open source Chinese base models like Yi and Deepseek and ask them about Tianmen Square yourself and see, they don't have any special filtering.
- meowface 3y agoI suspect they will crack down on that within the next few years.
- tcmb 3y ago> Russian and Chinese consumers are less interested in English-language articles. Isn't it just one additional step to automatically translate them?
- safety1st 3y agoThe NYT's strongest argument for infringement is that OpenAI is reproducing their content verbatim (and to make matters worse, without attribution). IANAL but it seems super likely to me that this will be found to be infringing sooner or later. Do I really want to use a Chinese word processor that spits unattributed passages from the NYT into the articles I write? Once I publish that to my blog now I'm infringing and I can get sued too. Point is I don't see how output which complies with copyright law makes an LLM inferior. The argument applies equally to code, if your use of ChatGPT, OpenAI etc. today is extensive enough, who knows what copyrighted material you may have incorporated illegally into your codebase? Ignorance is not a legal defense for infringement. If anything it's a competitive advantage if someone develops a model which I can use without fear of infringement. Edit: To me this all parallels Uber and AirBnB in a big way. OpenAI is just another big tech company that knew they were going to break the law on a massive scale, and said look this is disruptive and we want to be first to market, so we'll just do it and litigate the consequences. I don't think the situation is that exotic. Being giant lawbreakers has not put Uber or AirBnB out of business yet.
- jterrys 3y ago>IANAL but it seems super likely to me that this will be found to be infringing sooner or later. It better. Copyright has essentially fucking ceased to exist in the eyes of AI people. Just because you have a shiny new toy doesn't mean the law suddenly stops applying to you. The internet does its best to route around laws and government but the more technologically up to date bureaucracy becomes, the faster it will catch up.
- safety1st 3y agoYeah I mean I'm not even really a fan of how copyright law works, but I don't see how you can just insert an "AI exemption." So OpenAI can infringe because they host an AI tool, but we humans can't? That would be ridiculous. Or is "I used AI when I created this" a defense against infringement? Also seems ridiculous. Why would we legally privilege machine creation of creative works over human creation in the first place? So I don't see what the credible AI-related copyright law reform is going to be yet. Which means that either OpenAI is allowed to be the only lawbreaker in the country (because rich and lawyers), or nobody is. I say prosecute 'em and tell them to make tools that follow the law.
- jimmydoe 3y agoAnother way to look at it is to consider being stolen part of business model. There are massive number of piracy content in China, but Hollywood are also making billions in the same time, and in fact China already surpassed NA as #1 market for Hollywood years ago [1]. NYT is obvious different than Disney, and may not be able to bend their knees far enough, but maybe there can be similar ways out of this. [1] https://www.theatlantic.com/culture/archive/2021/09/how-hollywood-sold-out-to-china/620021/ https://www.theatlantic.com/culture/archive/2021/09/how-holl...
- deleted 3y ago[deleted]
- perihelions 3y agoSciHub was an early warning, IMHO, that there's a strong risk of the first world fumbling the ball so badly with IP that tech ecosystems start growing in the third world instead. The dominant platform for distributing scientific journal papers is no longer Western. Maybe SciHub is economically inconsequential, but LLM's certainly are not! Imagine if California had banned Google spidering websites without consent, in the late 90's. On some backwards-looking, moralizing "intellectual property" theory, like the current one targeting LLM's. 2/3rd of modern Silicon Valley wouldn't exist today, and equivalent ecosystems would have instead grown up in, who knows where. Not-California. We're all stupidly rich and we have forgotten why we're rich in the first place.
- deleted 3y ago[deleted]
- Geisterde 3y agoI have faith in your ability to make it through these difficult times.
- matheusmoreira 3y ago> We’re in a “mutually assured destruction” situation now, but instead of bombs the weapon is information. We've always been in that situation. Computers made the copying, transmission and processing of information trivial since the day they were invented. They changed the world forever. It's the intellectual property industry that keeps denying reality since it's such an existential threat to them. They think they actually own those bits. They think they can own numbers. It's time to let go of such insane notions but they refuse to let it go.