6 ms·
All this data is taken from common crawl, a web crawler that clones the web. How can companies such as OpenAI use that data when no licenses can be identified
by angrais 3y ago
All this data is taken from common crawl, a web crawler that clones the web.
How can companies such as OpenAI use that data when no licenses can be identified for most web pages or the associated images? In other words, the user/owner of the content has not agreed for it to be used in these ways.
What's people's experiences with training models using such data?
Is there any way to identify if my photos appear in common crawl or are used in such datasets?
- flangola7 3y agoSame way Copilot uses GitHub code and Midjourney uses scraped art images.
- brookst 3y agoKeep on mind that humans may also be learning from your work and emulating aspects of it.
- kolinko 3y agoYou have robots.txt standard that you can use to exclude crawling.
- gleenn 3y agoI believe there is an argument to be made that it is a derivative work. If you read a copyrighted book verbatim, that's an infringement, but if you summarize a book, that's fair-use. I don't think almost any of this has gone through the courts so there is a lot of legal peril and speculation though, obviously this is all super hot and pushing the boundaries so I'm sure we'll get clarification sooner rather than later.
- JimDabell 3y agoCopyright controls the right to copy, not any and all use. Just because somebody holds the copyright on something, it doesn’t mean they can dictate how it is used. “The copyright holder has not agreed for it to be used in these ways” is irrelevant. “The copyright holder has not agreed for it to be copied in these ways” is what matters. Analysing an image and adjusting weights in relation to that image is not making a copy. The images aren’t being copied into the model. You could argue that downloading the images in the first place to make that analysis constitutes copying, but these kinds of incidental copying aren’t normally considered within the bounds of copyright. If they were, you’d be committing copyright infringement every time you surfed the web.
- galaxyLogic 3y agoYou make a good point. I can browse the web and while I do I see pictures which get "copied" into my memory, somehow. However somebody looking at my brain with a microscope probably could not see those images in my brain, they are not localized that way. So copying something into my brain is not copyright infringement because I don't copy the image, I only allows the image to have some kind of affect on my brain. That is not copying I would argue. it is "experiencing". AI is a big brain which consumes the web almost like humans do. It "sees" the pictures on the web when it adds their characteristics to its associative memory. The AI then generates images. But unless it comes up with a clear copy of somebody else's work, it is not copyright infringement. I am not a lawyer so take this with a grain of salt.
- regularfry 3y agoCopying the image into your brain isn't what's material here. It's copying it into your browser over the network.
- robertlagrant 3y agoI don't think this is relevant.
- galaxyLogic 3y ago
- Gelotoooti 3y agoI also get inspiration from the internet. Am I not allowed to do this either? I download an image and save it and than use it. Either as inspiration in a collage or as inspiration when drawing something new. I don't think there have been a lot of people out there who really invented something uniquely new for a while.