3 ms·
I've taken courses from people in a few of these overlapping but historically-different camps recently, e.g.: - Frequentist statisticians - Bayesian statistici
by mjw 14y ago
I've taken courses from people in a few of these overlapping but historically-different camps recently, e.g.:
- Frequentist statisticians
- Bayesian statisticians
- Old-school AI researchers
- Statistical learning theorists
- Bayesian machine learners
- Engineers working on optimisation with noisy data
- Information retrieval folks
- ...
I'm really keen to see these guys starting to talk to each other and unify more of what they're doing around statistics (Bayesian and frequentist, parametric and non-parametric, generative and discriminative) as the common language and framework. Hopefully expanding the horizons of statistics a bit as a field in the process.
I imagine it'll take a while longer though, and some of the differences in terminology and talking-past-each-other can be a bit maddening in the mean time for those learning. What would be really nice would be a course offering a broad, well-rounded introduction to the various different philosophies to modelling data, their histories, interactions and overlaps, differing goals, strengths and weaknesses. It can be hard to get a sense of this when most introductory courses are taught by someone who's implicitly from one camp or another, even if (as is usually the case) they're not overly ideological about it.
One criticism of (some, not all!) statisticians is that they can seem to have a strange and rather limiting fear of computation which leads them to give undue preference to computationally simple models, even when the dataset isn't big enough to make compute time an issue.
I can see an argument for pedagogical reasons why one would rather not teach (or learn!) the details of fiddly optimisation algorithms -- but having a basic literacy in optimisation can free you up to treat the algorithm for optimising your objective as a black box to some extent. Feeling less "guilty" about this (whether the compute time itself, or remembering all the details of the optimisation algorithm) can be quite freeing in allowing you to think about the modelling itself in a more powerful, modular framework. This seems to be where machine learning has gotten a big advantage.
On the other hand machine learners can be frustrating in the way they reinvent statistical terminology and methods. Also in a gratuitous tendency to skip the modelling stage and go straight to inventing different objective functions to optimise, or even straight to the algorithms. Leading to opaque black-box methods which (due to the lack of probabilistic motivation) make it harder to reason about uncertainty in a principled way.
With the increasing popularity of Bayesian machine learning I think this is less of an issue though, it's bridging the gap between the two camps. One can also find a lot of more modern ML research using Statistical Learning Theory as a nice framework for principled frequentist analysis of their models.