3 ms·
Thank you for pointing out some bad phrasing on my part. When I said "small dataset", I should have said "small subset of all collected points that are initial
by cscheid 7y ago
Thank you for pointing out some bad phrasing on my part.
When I said "small dataset", I should have said "small subset of all collected points that are initially fed to the model". The issue isn't that you have to collect the data little by little. The issue is that, once you've given a linear model an input that is outside the distribution which you're hoping to model, nothing about the model can be trusted.
So you collect all the data points, but you don't give them all to the model at once. (I'm describing the RANSAC method here now) You start with a large number of "candidate models" that are all fit with a small number of input points, and then test which of the candidate models predict well the points you have not yet given the model. Then you feed the best of these candidate models only the points which it predicts well, and create a more refined, still outlier-free model. This can be proven to work in the presence of a small number of out-of-distribution points.
- derefr 7y agoVery interesting. It sounds to me like this approach exists outside any kind of Bayesian learning framework; i.e. an individual Bayesian agent couldn’t be expected to converge to a correct model here. Is that true? I would hazard that the RANSAC method sounds a lot like what you’d get out of a larger hybrid model, where e.g. many Bayesian agents with different priors are bred under a genetic algorithm after being ranked by their predictive power. (Like humans surviving to reproduce and pass down models to their children!)
- cscheid 7y agoAs far as i understand it the RANSAC arguments are very much not Bayesian (they're all about "if you repeat this an infinite number of times under an infinite number of new samples, then...", which is almost caricaturely frequentist). Still, you just gave me reason to mention one of my favorite "no, that won't work either" paper :) on how Bayes will not save you in the presence of model misspecification. Instead of butchering it any further, I'll just point you to this piece explaining the work, written by the author of the paper himself: http://bactra.org/weblog/601.html http://bactra.org/weblog/601.html