4 ms·
Some of those graphs are rather hard to pass, as the same domain is presented twice side-by-side, but with different X- and Y-axes. For example, comparing NN an
by timlod 3y ago
Some of those graphs are rather hard to pass, as the same domain is presented twice side-by-side, but with different X- and Y-axes. For example, comparing NN and linear regression models we see orders of magnitude difference on the Y-axis. Would really take the point of 'no change in RV model' home if they had the same scale.
Apart from that, many years ago, at a previous job, I prototyped automatic recognition of model degradation. I think one of the approaches I tried was comparing KL divergence between predictions on the test set and new data. This way you don't need to have the ground truth available (in which case you can make measures based on the residuals). Worked quite well iirc to signal when one should retrain and/or investigate context drift. Nowadays, with relative maturity in explainability methods, you can go much further and can much better find out what exactly causes the drift.
In general, I'm not a big fan of just saying 'error's up, let's retrain'. Why did the error go up? Understanding the data here better is usually quite valuable in the business context.