4 ms·
Taking data without consent is a real issue. There is still lots of data out there that is free of copyright. I'd be curious to see a model that is trained sole
by foruhar 4y ago
Taking data without consent is a real issue. There is still lots of data out there that is free of copyright. I'd be curious to see a model that is trained solely on public domain data (perhaps with an option to include creative commons-compliant data). I think there is plenty of knowledge that is in the free and clear to make a very useful LL and/or stable diffusion model. We may miss out on Wegovy and air fryer reviews, articles on the how to beat the stock market with Numpy, and manga art styles yet there is plenty of a few decades ago that would make for a useful "AI." Even Steamboat Willie may soon be in play.
- anileated 4y agoYeah, dated content would be the only reliable training data.
- pixl97 4y agoEh, you're just switching problems with the 'consent' model. I'm very much in the camp that the corpus of human knowledge is not some companies IP, this just pushes ownership further into hands of large and well monied companies and further baits patent/IP trolls to lock up collective knowledge.
- nightski 4y agoI look at it like it's some companies IP in the way oil/gas companies sell earth's resources. It takes a lot of work to transform raw crude into usable product, similarly OpenAI and others put a ton of money/resources/work into transforming that knowledge into a workable model.
- freejazz 4y agoGreat point, no issues with oil and gas companies as a business model
- pixl97 4y agoAgain this gets particularly messy. With oil there is very strong chain of possession. I can't copy your raw oil at little to no cost, and for the most part the next barrel of oil I pump out of the ground is not made of pieces of the past barrel of oil I pumped out of the ground. Each barrel of oil is a wholly separate entity. If I make all past oil disappear, you still have your barrel. Information is not like that at all. It is far more often a continuum of large bits of the past with small changes that redefine it's usage. If I took all bits of past knowledge out of your IP set, you'd be left with something useless in incomplete in almost every case. Trying to treat IP like a physical artifact leads to a multitude of failures.
- anileated 4y ago1. In your analogy a company is processing resources that humans didn’t create. Dead biomass from ages ago is not a fruit of your work, though note that even then you would expect extracting those resources from the ground to be taxed and proceeds to benefit you if you happen to be living on said ground. 2. Unlike petrol vs. raw oil, LLM output is not necessarily “better” than its source material. Indeed there’re plenty of authors that did both extensive work and eloquently written about it, so when LLM is asked a question on the topic their work is among the very few sources—when I am talking about LLM attributing output, I mostly mean instances like these (not when LLM aggregates some really common knowledge). The danger I described is when people are no longer motivated to do such work in the open, assuming it’s then scraped and monetized by LLM operators—or worse, undetectably modified by LLM designers to inject sentiment the original author didn’t subscribe to.
- nightski 4y agoMy point is that it's not just a matter of using the source IP directly. A lot of productive work goes into the creation of the model itself. Those weights & biases did not appear on their own. It could not be created without the source IP, but that doesn't mean the source IP is all you need to produce it. You need significant amounts of computing and human resources along with cutting edge research to produce it as well. While the art may be derivative in some cases, the model itself is unique and the value produced by these companies.
- anileated 4y agoTransforming oil into petroleum also involves a lot of productive work and it doesn’t go unrewarded and in fact has to be taxed appropriately. But generally in the case I described (topics in which just a few authors did most of the original work) I am not seeing any fundamental benefit if you compare this to a really good search engine—and such a search engine would benefit open information sharing, as it doesn’t do IP laundering.
- nightski 4y agoI guess I don't agree with the assessment that it is a glorified search engine at all. It's a lot more novel than that. I also agree that consent should be received before using an artists images in the training process. That said, if one could compute a training data image's contribution to the end result of a particular query it is entirely possible we could see a portion profits flow back to these artists from the use of their IP in the training process. But at the end of the day, when you train on billions of images, the end profit might be pretty minuscule. Any single artist's contributions might not actually matter all that much in the grand scheme of things. It's the combination of millions of artists that produce a result.