3 ms·
I'll drop my idea here since I won't be participating. Trade disk for privacy. Basically you keep all your initialized weights, gradient updates from the traini
by iandanforth 3y ago
I'll drop my idea here since I won't be participating. Trade disk for privacy. Basically you keep all your initialized weights, gradient updates from the training run and an index of what samples appeared in what batches. Then when you need to delete a sample, you find all the batches containing the target, reconstitute batches without that sample, and save those updated batches. You then take the initial weights and apply all the gradients that didn't come from contaminated batches. Finally you run a small additional bit of training with the cleaned batches.
This idea doesn't fully remove the influence of the target data (any previously saved gradient update from after a contaminated batch contains some information about the state of the network prior to update) but it may be a sufficient and efficient way to quickly reconstitute a network with far less influence from the problematic data.
Just an idea and I haven't tried it, so maybe it's bunk, but there you are!
- Filligree 3y agoDid you do any estimate of how much storage is required? On the face of it, I would expect the gradients to take about as much space as the weights. So you’d be checkpointing your network at every batch, in effect.