11 ms·
How to fix “AI’s original sin”
- mistrial9 2y agothis article hits many nuanced points well.. and it steers towards a future I can support as a WEIRD (common acronym for Westerners with a higher education?). too many detailed parts to respond to all of them in a brief reply.. I don't think these recommendations will apply to many parts of the non-West world.. here in the USA, this content makes a lot of sense to me
- rickydroll 2y agoThis is a good article, and I like the nuance it conveys. My takeaway from the article is that it highlights the need to identify what sources of information go into making a response. If a response uses the result of training on 10,000 documents and costs $0.0001, then every document gets its 1/10,000 fraction of 1/100 of a penny. When there is a payout scheme with information source traceability, I'm willing to bet that no-charge information will be used in preference to everything else. Alternatively, we could stop the copyright thuggery, treat training data as a worldwide Commons resource, and distribute the profits from LLM companies to sovereign funds.
- hammock 2y ago> When there is a payout scheme with information source traceability, I'm willing to bet that no-charge information will be used in preference to everything else. This did not happen in music.
- lmm 2y agoIt's gradually happening. Look at how many shops and restaurants today are playing cheap covers of recent hits rather than those recent hits themselves.
- muzani 2y agoIs this really the case? I'm a big fan of covers, and from what I've heard, nearly all of them share revenue with the original copyright holder. Except Hotel California - nobody is allowed to cover anything by the Eagles.
- pavon 2y agoThe Eagles have no say in the matter. Anyone in the US can legally record or perform cover songs as long as they pay government-defined statutory licensing fees.
- musicale 2y agoCover versions still require public performance licenses. Restaurants and bars typically pay for licenses from organizations like ASCAP and BMI, who distribute royalties to songwriters (rather than artists or record companies). Those licenses can cover recorded music playback and/or live music (usually with higher fees.) However, streaming and satellite radio stations have their own license regimes, which may include royalties to artists and to record companies (for streaming at least). So a shift to cover versions (perhaps owned by the satellite or streaming network itself) could lower payments by the network, so that could provide motivation for cover versions.
- lmm 2y agoRight. I think it's probably Spotify or equivalent somehow nudging them onto playlists of cover versions that are cheaper for them rather than the shops and restaurants explicitly making the choice. But it's a trend I've noticed that I thought was interesting, and suggests that cost is having an effect at some point in the chain.
- rickydroll 2y ago> This did not happen in music. Copyright thuggery is the reason why
- apantel 2y agoOne point that stood out to me that I agree with is this: > It is certainly worth knowing what content has been ingested. Mandating transparency about the content and source of training datasets—the generative AI supply chain—would go a long way towards encouraging frank discussions between disputing parties. But focusing on examples of inadvertent resemblances to the training data misses the point. I think that’s a good idea because that at least opens the possibility for affected content creators to knock on the AI company’s door and demand parley.
- mensetmanusman 2y agoThis is the middle class fighting the middle class; the billionaires that own both of types of organizations and the AI capital will inevitably use it to increase productivity and capture all of the excess monetary gains while the middle class shrinks.
- kapperchino 2y agoBruh stop with the middle class bs, there’s only working class and the owner class.
- redwoolf 2y agoThe middle class is the designation for the part of the population the owning class has duped into thinking they are temporarily embarrassed millionaires.
- labster 2y agoI’m not sure if you’ve seen housing prices lately but most of the middle class are technically millionaires now, just of the “land rich, cash poor” variety.
- redwoolf 2y agoBack that "most" up with facts. I'm not sure how much we can trust Zestimates.
- vsuperpower2020 2y ago[flagged]
- redwoolf 2y agoYou missed the point. There is no “middle class.” There is the proletariat and bourgeoisie.
- arh68 2y ago> respect signals like subscription paywalls, the robots.txt file, the HTML “noindex” keyword, terms of service, and other means by which copyright holders signal their intentions. And if they disrespect those signals & terms, and lie, what then? > the copyrighted content has been ingested, but it is detected during the output phase as part of an overall content management pipeline. Yes, but how can it be detected reliably? Considering there is much to be gained in fooling us. (think parallel construction)
- o11c 2y agoThe linked article cited Youtube's Content ID, so ... clearly reliability isn't expected.
- protocolture 2y ago>Yes, but how can it be detected reliably? Considering there is much to be gained in fooling us. (think parallel construction) I was fooling around with a fan constructed addon training module for novelai. I had a blast, reconstructing a few different narratives and sort of melding them together was a lot of fun. I let the tool name the characters. And the names it came up with were better than halfway decent. So I kept them. Turns out that while novel ai makes some sort of best effort to remove copyrighted proper nouns, the additional training module had reinserted some. After it had selected the first, it immediately selected his brother for the next one. If I hadnt gotten suspicious and googled the names in depth, I might have spread the story around. It wasnt ever destined to be published but I could see people falling into the same trap.
- kragen 2y agoi think plausibly being able to use youtube video as training data was the major reason for google to buy youtube in the first place. i'd be very surprised if youtube terms of service actually prohibit google from doing this also, while a lot of tim's thoughts are excellent, i strongly disagree with this part > When someone reads a book, watches a video, or attends a live training, the copyright holder gets paid reading books, watching videos, and attending trainings or other performances are not rights reserved to copyright holders, and indeed the history of copyright law carefully and specifically excludes such activities from requiring copyright licenses. consequently copyright holders do not in fact get paid for them. the first sale doctrine means that, in the usa (where the nyt has filed their lawsuit), not only can copyright holders not charge people for reading books and watching videos, they can't even charge them from reselling used books and videos this is fundamental to the freedom of thought and inquiry that underlie liberal civilization; it's not a minor detail
- deelowe 2y agoI doubt using it as training data was specifically the goal, but Google has always believed more data = more profit over the long term. This is why Gmail launched with unlimited storage.
- Semaphor 2y agoWas it unlimited? I only remember it being a decently high number at the time, far higher than any other freemailer.
- deelowe 2y agoMaybe you're right. It's been a while. Either way, the philosophy was always to make money off the data somehow even if we didn't know how at the time.
- TheDudeMan 2y agoCorrect. 1GB.
- renewiltord 2y ago> When someone reads a book, watches a video, or attends a live training, the copyright holder gets paid Bullshit. I have given many a book to a friend and they have passed it on. If this statement were true, then O(n) payments would have been made.
- c1sc0 2y agoI sometimes worry that qualms about copyright & ethics will make us lose the machine learning arms race with China. If « The unreasonable effectiveness of data » still holds true then we are in big trouble.
- deleted 2y ago[deleted]
- snowwrestler 2y agoI wouldn’t worry about it; in the U.S. the people with the qualms are not the people who are actually building the machine learning systems.
- SV_BubbleTime 2y agoOf course they aren’t the customers. That doesn’t stop them from having a disproportionately loud voice in the matter. Once you really understand what ESG is, you will understand why every company gets a rainbow logo this month, why companies seem eager to push their actual customers aside, and why the constantly-offended are catered to instead of mocked. Look at every AI press release. They all use “safety” over and over. Without that dogwhistle, they wouldn’t be in line for investment funding.
- SV_BubbleTime 2y agoNo need to worry about it happening, it already has. In terms of diffusers alone: Stable Diffusion 3 came out recently and immediately fell flat on its face. Terrible anatomy and basically unusable for human forms unless you use so many “unsafe” negative words that it’s “safety” training doesn’t destroy the result. In the meantime… Chinese models Lumina, PixArt, and Hunyuan are all Chinese projects or heavily contributed by Chinese researchers and companies. These are rapidly gaining steam. They’re far less “safe” with far less lobotomization. We are already losing the AI race in this realm exactly because of an over-reaction to “safety”.
- musicale 2y ago
- m3kw9 2y agoWhat happened to the generating training data hype train?
- protocolture 2y agoI mean the way to fix it is to recognise that its absolutely cool to train AI on publicly available knowledge. Like its not a sin. Maybe it comes from growing up with google hoovering up the internet, file sharing becoming common place and image boards making sharing copyrighted photos as reaction images the done thing. But I already feel like I own the sum total of human knowledge. I dont recognise sony or disney or the US government as valid inheritors or controllers of information. Its mine. And if there's a set of tools that can chew on that data to make new or interesting or collated or curated data then that just makes more data that I own. I own your art. I own your code. I own your stack overflow answers. I own every film and tv show and book from human history. And if that common ownership isn't the end goal then I dont know what is. If free use doesnt currently cover these cases it should be extended to cover them.
- idle_zealot 2y agoYou're not going to find many copyright abolitionists here. Apparently the idea that culture and information is a public good and not a product to be bottled and sold is so unpopular that it won't even be stomached in order to defend HN's favorite tech companies.
- diputsmonro 2y agoInformation in the abstract, maybe. But it's entirely different to assume ownership of artwork other people produce, either for profit or personal expression. Maybe I would feel a little differently if all the AI art I see came from a place of earnest artistic expression, but everything is just some kind of scam or another. You aren't going to convince me that honest artists deserve to have their art stolen (often by name) and livelihoods destroyed so that crypto grifters can generate NFT assets a little easier and scummy companies can lay off their art departments. This is a tool that will be used to concentrate wealth and keep the little guy down, pure and simple. And I'm definitely not swayed by the argument from this thread's OP of "I just want to steal everything I've ever seen and not have to do any hard work to create something myself".
- CaptainFever 2y ago
- musicale 2y agoAI's original sin isn't copyright violation. It's training machines on human-produced output.
- deleted 2y ago[deleted]
- wrs 2y agoI’m having a hard time understanding Tim’s point here. He seems to have retrieval and generation confused, or combined, or something. Of course in the retrieval (R in RAG) you can do attribution of the source material: you’re just doing a search and bringing up a literal excerpt of something from your database. You know exactly where you got it. But then for generation (G) you hand that excerpt to the LLM, and the only reason the model can “understand” it is the vastly larger corpus of text you trained the model on, whose origin is, due to the very nature of LLMs, smeared out of existence. That is the controversial aspect that (AFAIK) has no technical solution at present. He seems to imply that because attributing the R part is easy, attributing the G part should be too, but those are completely independent problems. Not to mention that you don’t have to do a retrieval step to generate things in the first place; the LLM alone can do that. The part about doing a retrieval on the output to see if it’s similar to something else is at least technically possible, but he handwaves past the problem of what on earth you’re supposed to do if you find something. YouTube doesn’t do a great job (e.g., people getting copyright strikes from their own performances of public domain works) and it at least has an unambiguous set of things to search for.
- antihipocrat 2y agoMy read on it was that the terms of use of content services can prevent that content being used to train LLMs. If LLMs were trained using data from these services then it's possible that this will be challenged in court. Should the court rule in favour of the content services, then the LLMs may need to be re-trained (or likely negotiate compensation).
- CaptainFever 2y agoRelated case: https://en.m.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn https://en.m.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
- wrs 2y agoHm, I thought the overall goal was that you would train LLMs on that data, but the owners of the data would be compensated when output was generated that was influenced by it. Somehow we have to be able to train LLMs on high-quality information, without having the resulting generative capability destroy the economic support for creating that information in the first place.
- diputsmonro 2y ago> Meanwhile, the AI model developers, who have taken in massive amounts of capital, need to find a business model that will repay all that investment. ... > The extreme case is these companies are no longer allowed to use copyrighted material in building these chatbots. And that means they have to start from scratch. They have to rebuild everything they’ve built. So this is something that not only imperils what they have today, it imperils what they want to build in the future. ... > "the only practical way for these tools to exist is if they can be trained on massive amounts of data without having to license that data." The simple and correct answer then is that they shouldn't exist. I don't care how much money they spent building a plagiarism machine, the expense and desire to build it doesn't excuse it being a plagiarism machine. Rich VCs don't get to ignore the rules just because they think they found a neat hack to make them billionaires. And after the shenanigans with the OpenAI board, the altruistic "for the betterment of humanity" argument has been revealed to be a thin sham.
- exe34 2y agothe confetti is out of the cannon now.
- neilwilson 2y ago> When someone reads a book, watches a video, or attends a live training, the copyright holder gets paid That's what copyright holders would like - hence why they try to restrict the licences so much these days. It isn't the case. Many individuals can read, watch and listen to a work where only one payment has been made. But really all the AI bots are doing is reading, watching and listening to the content and remembering it in pretty much the same way as a person does. Is an artist who has listened to a bazillion blues numbers and then constructs a new song based upon what they have heard and ingested in violation of copyright? If not, then neither is AI. It's just a probability matrix, not a replica. If the AI people have paid the correct fee to listen, watch or read the material and then sell what they remember it is no different from any trained professional. The contract has been fulfilled. The problem we have is that copyright holders have been dining off past glories for too long, and have gained the power to embed that rent for way longer than is sensible. What AI does is make those copyrights rot faster as it has a better memory than most people. That's good for society, because it forces copyright holders to produce more new stuff if they want to maintain their income. A rebalancing away from rent seekers towards regular producers would be good for everybody.
- nicbou 2y agoIn your last paragraph, are artists rent seekers and large companies that train AI producers? This is not how I feel when Google inserts its AI between my website and its readers.
- CaptainFever 2y agoWhat if we flip the dynamics: a hobbyist training a LoRA on Disney movies? This is just an appeal to... well, "big = bad". I believe what GP meant by rent-seekers are copyright holders who expect to own information for a hundred years, including all derivatives of it. It could be Disney, it could be an "original character do not steal" artist, and it could be "don't train on our AI" OpenAI. While regular producers would be those who use free licenses, or at least are OK with derivatives.
- 2y ago
- numpad0 2y agoThese[0][1] tweets showed up to my timeline recently. I don't know it's just an anti-AI luddite propaganda or not, and I do find people who do definitely not fit below caricaturization, but it seem to resonate with my sentiments and geostationary orbit gist of the matter: David Holz mentioned in today’s Midjourney office hours that they have more customers over 45 than under 18. He said you’re more likely to run into a 65-year-old woman than a teenager in MidJourney’s community.I think that’s a really interesting point about generative AI. It’s a bunch of old people who are telling you that it’s some mind-blowing invention. Young people are largely uninterested. They think it’s boomer art, not the guaranteed technology of the future. It’s only cool to olds. The internet was never like that—old people haven’t traditionally driven the culture. Tiktok and YouTube are often associated with younger people. Popular music and movies too. AI seems to only appeal to people who are past their prime. It lacks the It factor for people under 40. 0: https://www.threads.net/@thebrianpenny/post/C8aTGjUycs-/ 1: https://twitter.com/chiefluddite/status/1803704263148970255