5 ms·
> I think the worst possible outcome is a licensing regime that means that Disney or Paramount or Elsevier or whoever all get to have a monopoly on training lar
by oldkinglog 2y ago
> I think the worst possible outcome is a licensing regime that means that Disney or Paramount or Elsevier or whoever all get to have a monopoly on training large models within their niche.
Why is this the worst possible outcome? Companies using AI would be training with properties they own or have licensed appropriately, rather than the existing scheme of ignoring copyright law to extract $$$ from the creative works of ordinary people.
- mylastattempt 2y agoIn reality it would be a situation where a/the few big players control all the usefull datasets and trained models, and everybody has to pay in perpetuity (SaaS style) to use those, rather than being able to train their own. The copyright law being ignored will not be "solved" by pushing all the benefits to a few large players, but you will create new problems up ahead.
- rcxdude 2y agoBecause it would simply perpetuate the current copyright system which primarily rewards the large distribution companies disproportionately to the creators, especially smaller ones, while doing little to alleviate the effects on the job market for said creatives, in fact enhancing this effect by effectively feeding even more of the value into existing license holders as opposed to those creating new works (now instead of new works with existing franchises primarily rewarding those who own the franchises, this would effectively apply to all new works, if generative AI were to become mandatory to compete).
- dnissley 2y agoThey are extracting slivers of pennies on a per-work basis. Not $$$.
- bumby 2y agoThat’s still value though. It was illegal in the plot of Superman 3, I don’t see why it would be ignored here. Jaron Lannier has proposed a micro financing approach where the data owners are paid for their contributions to the data that is monetized by tech.
- xp84 2y agoOne nitpick with this: “Ordinary” people mostly don’t own any IP that earns them any income. Most who earn their income from IP are society’s elites. Perhaps the lowest-paid least elite IP profession I can think of is journalists, and I’m pretty sure they are doing works for hire and not owning the copyright. Arguably when it comes to the IP questions, ordinary people are far more likely to be benefiting from AI (for instance getting it to make art they couldn’t express because of lack of talent) than they are to be be robbed of their IP in a way that matters. Obviously the question of AI harming ordinary people in terms of automating mundane jobs is a very real one, but interestingly it’s totally irrelevant to the IP issues.
- bumby 2y agoI think the counter argument is that it works provide a pathway for “ordinary” people to monetize IP. Copyrights exist as soon as a creative work is made, it’s an entirely different IP tool than, say, a patent that has a more arduous process.
- visarga 2y agoYes but most LLM usage is for a single reader - the user who prompted the model. The output is rarely published. Maybe there should be a different rule for using LLM text to author public books and articles.
- bumby 2y agoAFAIK copyright laws do not make such a distinction. The most common exemption is fair use, but it requires the addition of some new creative aspect. Whether or not LLMs are creative or derivative may be an area that the courts are forced to define. Given the recent rulings that put this determination squarely on the court, I’m not sure the courts have the expertise to do so.
- visarga 2y ago> but it requires the addition of some new creative aspect That is one thing I can't get. Why is everyone not seeing the elephant in the room? The user and the user prompt - they add new intention and purpose to the material being referenced. Not only that, but a LLM response is usually shorter than a full length article or book. On top of that, LLMs retrieve from multiple sources when they use search, and respond from even more sources when they generate answers in closed-book mode. They are clearly doing synthesis. LLMs are not even good tools for infringement. You have to work really hard to put the LLM into an infringing mood, and it requires snippets from the target content, and only works like 1% of the time. If you have snippets, you probably have the whole thing already. Copying is so much easier and faster. LLMs would just reproduce approximatively a desired content. What a LLM gets from a source text in closed-book mode is on the same level with a image thumbnail from a full resolution image. What LLMs do well is to recombine ideas in new ways, as demanded by the prompt. Should that be infringing? If the answer is yes, then humans should also be barred from recombining ideas. After all, any human could secretly be using LLMs. But that would just tank creativity for fear of litigation. Isn't it strange how copyright promotes creativity by restricting it? Even the name "copyright" refers to copying, as it was initially envisioned. Now authors want to expand it to idea ownership. I think it's a power grab.
- chgs 2y agoIf they want to ignore copyright great, let’s change the law so copyright only lasts say 10 years. But these companies want to have their cake and eat it.
- Kim_Bruning 2y agoNot sure why Paramount is on that list, but ofc. Disney and Elsevier are in the business of applying copyright law to extract $$$ from the creative works of ordinary people. (Disney: takes fairy tales from the public domain. Elsevier: You pay them to get published)
- pas 2y agoWhile I don't clap for Disney at all, it's important to note that the fairy tales are still there where Disney found them. The problem with Disney is their extremely aggressive and repeated push to safeguards their own golden goose. (And IMHO they have a serious quality problem. "Somehow Palpatine returned.") In general the quasi-forever current copyright regime is simply too blunt, and arguably waay over the optimum point in terms of length of granted protection with regards to "incentivizing the creation of arts", and of course it's incentivizing the collection of money-making properties and flooding the market to extinguish most of anything else. (Naturally attention is a limited resource, so "arts is a zero sum game".)
- viraptor 2y ago> would be training with properties they own or have licensed appropriately Cool - you want to train on medical data? That will be $1M per paper per day to Elsevier for a licence. (Apply similar for movies, news, books, etc.) Ensuring no knowledge remains "common". The moment they can, the big data holders will charge whatever they can get away with. For an average person that may be even worse than the current content stealing.
- pas 2y agoPapers are ridiculously low signal-to-noise when it comes to knowledge. (see https://markusstrasser.org/extracting-knowledge-from-literature.html https://markusstrasser.org/extracting-knowledge-from-literat... ... but sure, LLMs might have a better extraction rate) And, importantly if papers become so valuable ... authors will suddenly start to want their cut and will go to the publisher that pays the more. And so on. It would be amazing if papers would worth that much (as it would mean they can help creating at least that much value downstream), and it would mean there would be a lot of money in writing new papers. (Oh imagine that flood of low quality shit! Oh no, it's almost the same as now! Ehhh.)