4 ms·
>It's the robbery of all of our culture to sell it back to us at a mark-up Would regulation help with that? Right now you can download free models that have be
by pingou 15d ago
>It's the robbery of all of our culture to sell it back to us at a mark-up
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free.
I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".
- embedding-shape 15d ago> With regulation and compensation, only rich companies would be able to do that Well, with some imagination, you can have regulation that forces companies to open up, not just close down. Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download. Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
- ben_w 15d agoI am unclear how this would help anyone? Any argument that writers and artists lose from these existing, would remain unchanged.
- embedding-shape 15d agoThe argument was "It's the robbery of all of our culture to sell it back to us at a mark-up". Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from. Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content. I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.
- CJefferson 15d agoWe don’t have to treat people reading books and companies stealing all human knowledge the same. Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
- hlynurd 15d ago>companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves really not the same entities here
- CJefferson 15d agoNo, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.
- hlynurd 15d agoYeah that's more fair
- cgio 15d agoNot at leaf level, but if you trace the trunk, pretty sure you end up on the same one.
- ipython 15d agoWell, it's kinda converging, because Napster and Microsoft have teamed up to build a multimodal interactive video agent through a simple proxy API (this is a direct quote from Napster's blog post) https://www.napster.com/blog/napster-heads-to-microsoft-build-with-omniagent-api https://www.napster.com/blog/napster-heads-to-microsoft-buil...
- torginus 15d agoMorally speaking, one of the issues of modern society is the idea that knowledge should be free which was partially started by the file sharing movement, which didn't really move society in the right direction imo. Free means the same as worthless, which inherently isn't true - since information takes time to consume in some form, and your time isn't worthless. Therefore even if you could listen to all songs theoretically for free, you would need to spend an inordinate time doing that. When I was a kid, getting a CD from your favourite band was a major expense, getting a video game even more so. But it formed a sort of emotional attachment (and not even just for me), my friends talked about how 'band X''s new album was amazing or a stinker. Since there were multiple bands making similar kinds of music, choosing to be a fan of one but not the other carried real monetary weight. Nowadays you just fish out a song you think you would like out of the endless sea of Spotify, no different from prompting an LLM. No, Spotify didn't make me enjoy music more. Same applies for Steam & videogames. Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it. I'm sure most musicians would trade a million fans who might check them out if they're in town, to ones who think their music changed their perspective in life. Also the process of creating a song that vaguely appeals to millions is different from making one that speaks to a thousand. This is a fundamental issue of modern capitalistic society, similar to the Marxist idea of 'alienation' - once something is cheap to get, you don't appreciate the effort that went into making it. And if your customers don't care about the thing they get, producers won't make an effor to make it good either. And once nobody cares, people even forget what a quality product is like.
- mitxela 15d agoCulture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.
- loloquwowndueo 15d agoThe what thread, now? Got a link instead?
- n_plus_1_acc 15d agohttps://news.ycombinator.com/item?id=49737787 https://news.ycombinator.com/item?id=49737787
- steveBK123 15d agoWell theres at least two different buckets of this. First is the scraping of the open internet. The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another. Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens. The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
- robinsonb5 15d agoAnd yet a third bucket is the license-laundering of GPL code when the entire github corpus was vacuumed up.
- vitorfblima 15d agoCopy one book, and you're a thief. Copy thousands, and you're a VC.
- deleted 15d ago[deleted]
- fzeroracer 15d ago> Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
- Luker88 15d ago> and they would definitely not give it back for free. ...not like they are doing it for free now either. open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business. > I put "stolen" in quotation marks because it's still unclear if we can call that stealing It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement. This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright. They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail. So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot. All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up. Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval. Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
- barnabee 15d agoRegulation that said something like “we own 50% of your profit or 20% of your revenue, whichever is the larger” would. If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
- jappgar 15d agoRegulation can mean all sorts of things, including declaring the models themselves illegal. Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.
- kshri24 15d ago> Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing". It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways. We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified. EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models. So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
- echoangle 15d agoThe multiplication comparison makes no sense because humans don’t learn by multiplying numbers to change weights. It’s like comparing the lubrication oil consumption of a car to the cooking oil consumption of a human to compare the carrying capacity. That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn. Otherwise, a human would learn much less in their whole life than a GPU does in one second.
- Dead_Lemon 15d agoBreak up the multi trillion, multifaceted companies into the their separate facets. These companies all thrived on far significant smaller portfolios in the past. Cap the size a company can grow. Regulate the amount of compute they are allowed to use.
- 2frrrr 15d agoBut china!
- ElProlactin 15d agoNone of this actually stops the profiteering, or reverses the harm already done.
- dofm 15d agoLegally mandating that companies open the weights of, say, 18-month-old models might change their minds a little. I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.
- the_other 15d agoWhen a human learns, they carry that forward into their future ventures. They might use their learning to recreate the original work and profit from it without attribution. In the West we view that badly and have legislation to protect from some abuses. But the same person might later collaborate with the original author, or make a derrivaive work that improves on the original (a la most science). AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two: - either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing - or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)
- tripzilch 15d ago> Would regulation help with that? I mean it's already happened, right? I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...
- saynay 15d agoA simple law stating that laundering data through an LLM does not constitute "fair use", that the outputs can be subject to copyright claims of the original authors, and that the outputs themselves do not qualify for copyright protections would go a long way. It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.
- mina_bridge 15d ago[flagged]