5 ms·
> the synthesized images would clearly be derivative works, and thus copyright holders have rights to the (portion of the) output which is derivative. Good luc
by ttt0 5y ago
> the synthesized images would clearly be derivative works, and thus copyright holders have rights to the (portion of the) output which is derivative.
Good luck proving it. So in practice it's basically legal.
- cwkoss 5y agoIf they find their image in the training dataset via discovery, that would prove it. I think the trickier part is figuring out that your image was used in the dataset or convincing the judge that you have standing to compel that discovery search. If you have a bunch of get requests from an AI company's corp IP that could do it.
- ttt0 5y agoI'm not familiar with ML, but I assume you can simply get rid of training dataset after the fact, right? You don't have to keep it around forever? > trickier part is figuring out that your image was used in the dataset or convincing the judge that you have standing to compel that discovery search Yes, that's my point.
- cwkoss 5y agoOnce the model is working, the training set can be deleted - but it would be hard to modify the model in the future without. I believe there are ways to add to a training set without retraining on the whole thing, but it's trickier and the math is way over my head. I think in many cases people just retrain over the whole corpus if they want to iterate on the model - so most commercial ML models probably hold on to their training datasets and consider it a business asset.
- Mehdi2277 5y agoIt's very doable to add without it and my experience is I've often worked with models trained online where the data has a retention of a few months so any old enough model will have been trained on data points that don't exist anymore. Large tech companies tend to not keep all data for a mix of storage cost, lower value for some domains (for recommendations freshness is quite useful), and various privacy concerns. Privacy alone means you should have ways to delete data entirely for many problems. If you train with user data and the user requests to delete all info on them through gdpr that includes your training data. I can't remember the amount of time gdpr allows for deletion (I think 30 days or maybe 60), but if you keep a lot of your data with a lower retention than that by default that makes deletion a lot easier. There is other data that has naturally very long retentions (sales orders) where even with deletion requests some aspects should be kept just for financial compliance.
- ohazi 5y agoThere are ways to hornswoggle a trained machine learning model into generating new outputs that "kinda sorta" look like bits and pieces of an input that was used for training. It's tricky, and not always reliable, but it would be a reasonable thing to attempt. This is why you shouldn't use sensitive material as training data. If you're a startup and you're using all of your customers' private data to train a single model that then gets used by all of your customers -- guess what, you're doing it wrong! It's the same reason you wouldn't teach your five year old how to read by letting them pour over you and your spouse's text message history. I mean, you could, but then you shouldn't be surprised when they inadvertently utter something embarrassing that they don't necessarily understand, but whose origin is obvious to everyone else at the party.
- bencollier49 5y agoOnce you've got your Neuralink in, they'll be able to do the same thing for anything you create. Moderately influenced by a piece of music you heard when you were three years old? Pay up!