10 ms·
Apple, Nvidia, Anthropic Used Swiped YouTube Videos to Train AI
- GaggiX 2y agoThis is about the Pile dataset, of course we don't if it has been used to train the commercial models we use or just for the research papers mentioned in the article.
- porphyra 2y agoI wonder how much of video generative AI depended on the open source project youtube-dl/yt-dlp.
- dcchambers 2y agoSadly stuff like this makes me think Google is going to work even hard at preventing tools like YT-DL from working. IMO data is Google's biggest moat in the AI race - and I suspect they'll do whatever they can to keep it.
- CuriouslyC 2y agoGoogle's moat is deep ranks of machine learning engineers and top flight SWE talent, coupled with probably the best infrastructure in the business. Their data moat is weak, the data they can actually train on is mostly public, and people trying to prevent scraping are always on the losing end. Meta is the one with the huge data moat.
- sidcool 2y agoIt's an open secret. The only thing missing is the clear evidence.
- middlefing 2y agoOP used swiped content to deliver this info.
- nerdjon 2y agoThat is not the same and I really feel like the distinction should be obvious. Most articles will link to sources from others and then build on top of them for their own article. Actually giving sources for your work. That isn't swiping the content to write an article. Yes there are bad actors in that regard, but most play by the accepted rules here. With the AI work, attribution of source is gone. Who did the work that you are benefiting from is gone. The people that benefit are those that made the AI and the people using it, skipping the source for the training data.
- nickpsecurity 2y agoI believe that the law allows us to link to others’ online content. The reason might be that the link takes you to their site where they control the presentation and can benefit from your viewing it.
- xhkkffbf 2y agoThey're not the only ones. I've seen other AI firms like Adept doing the same thing. I smell a class-action suit from all of the video makers who weren't protected by YouTube.
- verst 2y agoAnd furthermore, Google claims that their own training of models on YouTube data is permitted.
- danans 2y agoThe training of models seems less of an issue than what you do with them. IMO, the most problematic part is generating content in the style and voice of the original human creator of the copyrighted content. Unfortunately, this is also currently among the most user-attractive for generative AI trained on copyrighted content.
- threatofrain 2y agoFor anyone familiar with the legal landscape, except for scenarios where AI products make reproductions of their training material, why isn't this covered under fair use? Don't humans basically do the same thing when attempting to create new music — they derive a lifetime of inspiration from the works of others?
- vouaobrasil 2y agoBecause AI often replicates things much more closely than what fair use would constitute, and doesn't label sources like when you are quoting. And it's generally harmful for humanity, too.
- threatofrain 2y agoI'm specifically interested in situations where ML products do not do simple reproductions of copyrighted material. I'm aware that it's difficult to even know the space of output and to "align" the model correctly. Are we normally required to label sources when referencing other copyrighted materials, whether in songs or movies or otherwise?
- johnnyanmac 2y agoDepends on how much of the source you used (which is in line with how fair use works, unless you're developing parody). Given that AI is using the entire source: yes. As we know with scraping cases, the amount data and time also may play a role in determining fair use (think in terms of buffet ettiquite. "all you can eat" does not in fact mean "you can eat it all by yourself"). Funnily enough LinkedIn (owned by Microsoft) did argue successfully in court against scraping a website.
- stetrain 2y agoThere are lots of things that are not simple reproductions that are not fair use. If I take ten of your copywritten photographs and stack them on top of each other in Photoshop with transparency, the output is not a simple reproduction of your work. If I sold that for commercial purposes you would be upset with me and likely have a copyright case. That's an obvious example, but my point is there aren't super clean-cut definitions for these things, and it's not settled case law yet which side current AI training and content generation falls under.
- gnrlst 2y ago[flagged]
- deleted 2y ago[deleted]
- mmanfrin 2y agoI wonder how many of the creators complaining make react videos.
- luqtas 2y agowith click-bait titles like: Apple hacked ME to put MY videos on their AI! & a thumbnail with the logo of Apple at the top-right and their face doing a weird expression, occupying 75% of the .png with a single bright color background
- johnnyanmac 2y ago> Proof News also found material from YouTube megastars, including MrBeast (289 million subscribers, two videos taken for training), Marques Brownlee (19 million subscribers, seven videos taken), Jacksepticeye (nearly 31 million subscribers, 377 videos taken), and PewDiePie (111 million subscribers, 337 videos taken). Some of the material used to train AI also promoted conspiracies such as the “flat-Earth theory.” I'm not going to grok through all their thousands of videos, but these specific creators sure didn't get this big on reactions only.
- RcouF1uZ4gsC 2y agoHow many of the YouTube channels depend on fair use themselves? For example, Jacksepticeye is listed as having their videos used. Looking at the channel, it seems like a lot of it is is recordings of them playing video games. Is the company that produced these games being compensated?
- danans 2y ago> How many of the YouTube channels depend on fair use themselves? IANAL, but recording portions of copyrighted content or using excerpts thereof is covered under fair use. It is not yet known whether reproducing copyrighted content in substance or style using generative AI is covered under fair use.
- zerocrates 2y agoIn this specific video game context, fair use is almost totally up in the air. It's kind of crazy given how much of a market there is on videos and streaming of games, but it's how it is. Both "sides" of that argument prefer the status quo to the risk that a legal decision would go against them. There was a period of time several years ago when some publishers were not allowing some kinds of streaming, and there was a need for them to post public statements allowing things like Let's Play videos. Even now you'll have the occasional game where the publisher just says streaming isn't allowed, or is only allowed under restrictive terms.
- danans 2y agoNow imagine combining both video game streaming and generative AI to create generated video game streams. Given the relative simplicity of video game worlds, it should be far easier to generate those than than photorealistic video (i.e. Deep mind Veo, OpenAI Sora). Yes, it might just saturate the world with low quality content, leaving the good stuff still distinguishable, but many content business models are built on low quality content.
- nerdjon 2y agoYou know what game is it though. You know where they are getting their sources from. As much as I hate react videos, at least the situation is the same. You know where the source is from and can go to the original if you wish. Show me the attribution on your generated content for any of these creators.
- trumps-ear-bug 2y ago[flagged]
- johnnyanmac 2y agoShocking or not, I think "this company is potentially breaking the law" is always worth calling out.
- raviparikh 2y agoWhether covered under fair use or not, the laws around copyright today did not anticipate this use case. Congress should pass laws that clarify how data is and isn’t allowed to be used in training AI models, how creators should or shouldn’t be compensated, etc - rather than speculating whether this usage technically does or doesn’t comply with the law as-is.
- 2OEH8eoCRo0 2y agoHow would you ensure compliance?
- bjt 2y agoCreate a private right of action. If creator A can show that AI trainer B used their works (e.g. like how we've seen Getty watermarks show up in AI generated pics), then they can sue for $X dollars.
- johnnyanmac 2y agoI think what really sizzles me is that some of these same companies helped develop such strict enhancements to copyright to begin with in the realm of software. So I'm not falling for the crocodile tears when they get caught in the very snare they used to litigate thousands of other companies for and bully potentially millions more with just because now it's more profitable to tear it down. Made your bed... And yes, regardless of results I agree there should be new laws made. But we know Congress in the US this year has been a roller coaster, to put it lightly. And I don't even think this is top 5 of what congress needed to codify into law properly. So all the short term work will be the judicial branch interpreting what few laws we do have.
- tzs 2y ago> I think what really sizzles me is that some of these same companies helped develop such strict enhancements to copyright to begin with in the realm of software. What enhancements are you thinking of?
- deleted 2y ago[deleted]
- TIPSIO 2y agoI imagine data laundering (?) is common. E.g.: Nike needs to produce a large amount of clothes. They hire an oversea company who commits to the order. They set strict rules -- no child labor, certain quality controls, etc... This company then subcontracts anyway possible and delivers the order, gets paid, and dissolves. Messy but Nike's hands are clean. With AI, same thing but with videos and other forms of data. Hence why a question "did you train with Youtube?" to a certain CTO is so difficult to answer.
- matheusmoreira 2y ago> Hence why a question "did you train with Youtube?" to a certain CTO is so difficult to answer. Make it so they're automatically guilty if they can't provide a definitive negative answer.
- karolist 2y ago> “YouTube’s terms cover direct use of its platform, which is distinct from use of The Pile dataset. On the point about potential violations of YouTube’s terms of service, we’d have to refer you to The Pile authors.” so basically "we stole from a thief therefore we didn't steal" excuse?
- hollerith 2y ago"A thief gave it to us", you mean.
- deleted 2y ago[deleted]
- jmyeet 2y agoSo there is a clear effort to build enclosures around various of corpuses of material that could or would be useful to train AI. Thing is, people read books, they watch videos, they listen to music, they see and produce art and so on. How is training data distinct from "human training data"? One could say the quantity. We're currently dealing with statistical learning models that require a huge quantity of training data. This is temporary. At some point you will be able to train an ML system with less because humans can be trained with less. What then?
- blackeyeblitzar 2y agoI don’t see anything wrong with these companies using YouTube content to train AI in a sense. I think the creators of the videos should be fairly compensated and their permission should be sought, but I don’t think of Google/Alphabet in that way. Sorry but even if Google runs the YouTube platform, I just don’t think they ethically or morally have an exclusive right to the content the world creates, just because they have various monopolies that are immune to competition due to anti-competitive moves and the power of network effects. As far as I am concerned they are a utility service that needs to be heavily regulated.
- londons_explore 2y agoIt's awfully hard to imagine any specific kind of harm that video creators will suffer by having AI trained on their subtitles...
- ikekkdcjkfke 2y agoIt's the chicken egg problem (?). ChatGPT couldn't make LLM so valuable without stealing. It's just forcing the freemium model, if google accepts settlement/later payments without any crime charges
- MiguelHudnandez 2y agoLegal concerns aside, aren't Youtube captions primarily AI-generated in the first place? I know some authors meticulously hand-craft their captions but that can't be the case for the vast majority of videos. Therefore isn't training AI on this basically poisoning your own model? The caption quality is good but there are mistakes in pretty much every video I watch with captions.
- eigenvalue 2y agoThey are almost certainly extracting the audio and then using Whisper or other superior speech recognition models. I made a free tool which can do this very efficiently for whole playlists of YouTube videos, so I'm sure they can do the same: https://github.com/Dicklesworthstone/bulk_transcribe_youtube_videos_from_playlist https://github.com/Dicklesworthstone/bulk_transcribe_youtube...
- IronWolve 2y agoPretty sure, web scraping has been upheld as legal when microsoft lost its case with companies scraping linkedin. And generative content is also legal, thats even includes reposting a copyrighted video even if there is discussion video over the video. Thats an extreme case of fair use, but it shows a wide use case over a copyright video. Personally been using fabric ai tool, since it can summarize youtube videos, so I dont have to watch an hour+ video or read a very long article/journal, just gives me a summary, top talking points or even break it down for tech points. https://github.com/danielmiessler/fabric https://github.com/danielmiessler/fabric
- leereeves 2y ago> And generative content is also legal, thats even includes reposting a copyrighted video even if there is discussion video over the video. Citation please.