3 ms·
Can somebody explain to me the underlying theory of this type of book? The only books that have ever felt coherent to me start with p(data, unknown) as being a
by tristanz 11y ago
Can somebody explain to me the underlying theory of this type of book?
The only books that have ever felt coherent to me start with p(data, unknown) as being an approximate model of some domain. Everything then follows smoothly as inference, modeling, and computational methods or shortcuts.
- jey 11y agoI agree, but it's also nice to have a toolbox of actually tractable algorithms. Principled practical data science should approach the theoretical ideal (i.e. p(unknown|theta)), but often has to use some approximations that we actually know how to implement efficiently.
- jey 11y agoHeh, I meant s/theta/data/.
- hyperbovine 11y agoI used a previous version of this manuscript in a class, and it was more descriptively called 'Computer Science Theory for the Information Age'. It's a collection of stuff computer scientists ought to know about probability, (high dimensional) statistics, ML, randomized algorithms, etc. They seem to have sexed up the title a bit; at least it does not contain "Big".
- jules 11y agoThere are two approaches. One approach starts with the data and asks what algorithms can we find to infer something about the data, and the other approach starts with an algorithm and tries to find a dataset on which the algorithm infers something good. This is even more general than machine learning; in many cases you either have a problem and are searching for a solution, or you have a solution and you are searching for a problem. Often you need a mix of both (a bidirectional search if you will). It looks like this book is a mix of both. In the end though the only thing that matters is whether a particular procedure has predictive value. Even with the most principled Bayesian analysis the model or prior are strictly speaking wrong (usually both), because the real world is complicated, so even there you're only left with testing predictive performance in practice. Careful probabilistic modeling is only valuable insofar as it lets us focus our search for inference procedures on those that are more likely to work in practice. A good counterexample is deep learning, which does not (or did not) have a solid probabilistic justification, but works extremely well in practice.