3 ms·
I'm not familiar with ML, but I assume you can simply get rid of training dataset after the fact, right? You don't have to keep it around forever? > trickier p
by ttt0 5y ago
I'm not familiar with ML, but I assume you can simply get rid of training dataset after the fact, right? You don't have to keep it around forever?
> trickier part is figuring out that your image was used in the dataset or convincing the judge that you have standing to compel that discovery search
Yes, that's my point.
- cwkoss 5y agoOnce the model is working, the training set can be deleted - but it would be hard to modify the model in the future without. I believe there are ways to add to a training set without retraining on the whole thing, but it's trickier and the math is way over my head. I think in many cases people just retrain over the whole corpus if they want to iterate on the model - so most commercial ML models probably hold on to their training datasets and consider it a business asset.
- Mehdi2277 5y agoIt's very doable to add without it and my experience is I've often worked with models trained online where the data has a retention of a few months so any old enough model will have been trained on data points that don't exist anymore. Large tech companies tend to not keep all data for a mix of storage cost, lower value for some domains (for recommendations freshness is quite useful), and various privacy concerns. Privacy alone means you should have ways to delete data entirely for many problems. If you train with user data and the user requests to delete all info on them through gdpr that includes your training data. I can't remember the amount of time gdpr allows for deletion (I think 30 days or maybe 60), but if you keep a lot of your data with a lower retention than that by default that makes deletion a lot easier. There is other data that has naturally very long retentions (sales orders) where even with deletion requests some aspects should be kept just for financial compliance.
- ohazi 5y agoThere are ways to hornswoggle a trained machine learning model into generating new outputs that "kinda sorta" look like bits and pieces of an input that was used for training. It's tricky, and not always reliable, but it would be a reasonable thing to attempt. This is why you shouldn't use sensitive material as training data. If you're a startup and you're using all of your customers' private data to train a single model that then gets used by all of your customers -- guess what, you're doing it wrong! It's the same reason you wouldn't teach your five year old how to read by letting them pour over you and your spouse's text message history. I mean, you could, but then you shouldn't be surprised when they inadvertently utter something embarrassing that they don't necessarily understand, but whose origin is obvious to everyone else at the party.