3 ms·
Seems like a good collection of standard ML techniques, introduced with a fairly unified mathematical notation and quite a few proofs. Quite the Herculean effor
by c7b 3y ago
Seems like a good collection of standard ML techniques, introduced with a fairly unified mathematical notation and quite a few proofs. Quite the Herculean effort (600 pages!). It just seems to me like they're putting the emphasis on the stuff that is more straightforward to formalize rather than the stuff that would be interesting to understand.
Look eg at the SGD chapter. I picked this because I think optimization is one of the areas where mathematicians actually can and do make impactful contributions to ML. But then look at the chapter in the book: most of the proofs are fairly elementary (like bias-variance decompositions or Jensen inequalities), some more interesting theorems (on convergence) are cited from the literature and do not build on the lemmata, and the sub-chapters on the actually interesting methods like ADAM,... are completely free of proofs or theory. It seems to me that after reading the chapter, a reader will have a good understanding of modern SGD methods and how we got there, but they won't necessarily be much wiser about why those methods work, other than having a good intuition confirmed by numerical experiments. If that's the outcome, then I wonder what the fuss proving all the basic stuff was all for. Wouldn't it be more useful to dedicate the space to convergence proofs for ADAM (which do exist) rather than showing lots of stuff like E(XY) = E(X)E(Y) for independent random variables?
That's just one chapter, I may not be doing them full justice here, although I did read through a few others as well. I first got this impression from the ANN chapter, which is ripe with long proofs for rather basic and uninteresting stuff, and from the physics-informed neural networks paper (which I actually find really nice, although it suffers a bit from the same problem as the SGD chapter). I don't want to be too critical here, it is nice in general to move towards a more rigorous and unified exposition of ML methods, and their approach should extend to the more technical results as well, just questioning where they drew the line of what to include and what not.
- modeless 3y ago> they won't necessarily be much wiser about why those methods work, other than having a good intuition confirmed by numerical experiments. This is the state of the field as a whole, isn't it? > Wouldn't it be more useful to dedicate the space to convergence proofs for ADAM (which do exist) Convergence proofs don't really explain why Adam tends to work better than other methods. It's hard to blame them for not being able to explain things that, currently, nobody understands. But I guess it kind of undermines the idea of a theory-heavy approach to teaching if the theory we have can't predict the things that are actually important.
- go_elmo 3y ago*Cant predict the important things _yet_
- c7b 3y agoADAM is known to have better convergence bounds than other methods. Theoretical bounds may not explain the full story of why a method works well, but it is how mathematicians reason about it. I'm only blaming them for not sharing those relevant parts of what we already know. My bigger pain point even is how they choose to allocate their space: the theorem statements for the most relevant results are missing, the proofs for the more interesting theorems are just citations, while the proofs for basic and arguably tangentially relevant lemmata from eg probability theory take up pages and pages.