4 ms·
Can you expand on this? If you’re monitoring for data drift and retraining every so often after deployment (not just “set it and forget it”), what are the probl
by willj 7y ago
Can you expand on this? If you’re monitoring for data drift and retraining every so often after deployment (not just “set it and forget it”), what are the problems that can happen?
- wrkronmiller 7y agoWARNING: NOT AN ML EXPERT. I believe this falls under the category of "data snooping" wherein you are effectively creating a model-of-models and increasing your degrees of freedom. That increased level of complexity/number of DoF means you are far more likely to over-train. You are more likely to pick a model with no predictive power that "happened to be right" about past data.
- ldoughty 7y agoIf you're not careful, you're training your algorithm to continue promoting a bias... Perhaps one that you're not even aware of... Or perhaps someone else's bias imparts a bias on your algorithm (like turning your chat bot into a Nazi sympathizer -- an extreme example, but it could also be someone else's refusal of loans to minority groups, and your algorithm might look at debt ratio as a factor) You can't keep assuming your inputs are always good... and while it may be okay to do some automated retraining, you really need a skilled person to revisit the inputs/outputs...