6 ms·
I see a completely different attack vector here. Lawyers. If you are selling a product or service that has been trained on a dataset that contains copyrighted
by OldHand2018 5y ago
I see a completely different attack vector here.
Lawyers.
If you are selling a product or service that has been trained on a dataset that contains copyrighted photos you don't have permission to use and I can "prove it" enough to get you into court and into the discovery phase, you are screwed. I'll get an injunction that shuts you down while we talk about how much money you have to pay me. And lol, if any of those photos of faces was taken in Illinois, we're going to get the class-action lawyers involved, or bury you with a ton of individual suits from thousands of people.
That link at the bottom about a "safe harbor" you get from using old datasets from the Wild West is not going to fly when you start selling.
- laura_g 5y agoThis would be membership inference attacks - https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=7958568&casa_token=edBGtFGrJlwAAAAA:eLObvr85G6__HNWXaAWmdqcK-ljtTy_eniFBv4BM3EBojJ7pza3n-GxsnOmqC8GFjj5ZCzZb https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=7958568...
- OldHand2018 5y agoOh excellent. But of course the key addition is handing off this information to lawyers who use it to shut you down and/or extract money from you. If you are using some torrent of a dataset, nobody is indemnifying you, and once you get to the discovery phase of a lawsuit, they are going to know that you intentionally grabbed a dataset you knew you shouldn't have had access to. Treble damages!
- NoGravitas 5y agoI dunno, Microsoft seem to think they can get away with training autocomplete on copyrighted source code that they don't have permission to use.
- KingMachiavelli 5y agoIIRC simply training on copyrighted material is completely fine or at least you can claim fair use. As long as the market of the copyrighted material is not 'AI data training set' then it should be OK. Essentially scraping images from the internet is OK but using a pirated copyrighted commercial AI data training set is not. (Fair use doesn't necessarily exclude use for a commercial/sold product.) But if the AI model just spits out copyrighted material verbatim then that is still owned by the actual copyright holder.
- OldHand2018 5y agoThe article contains a link to an Adobe blog that talks about fair use of copyrighted material. It references the fair use doctrine in a way that is not fully analogous to this type of use and mentions the Google books case. It also mentions that this is not settled law. It's clear that the author wants it to be fair use, but that might cloud their analysis. Keep in mind that Google was scanning books that the legitimate owner of the physical books gave them permission to scan. If I buy a book and want to use it to train my model, fair use says I am free to do so. If I grab an unauthorized torrent of a training set, itself containing images were not legitimately purchased or licensed, there is absolutely no case law that I know of that says it is ok. I have to spend my money on lawyers trying to argue that I'm in the clear with no guarantee of success. Maybe I'm wrong - I'd love to hear a convincing argument to the contrary!