7 ms·
Those Wikipedia pages are kind of awful for pedagogy, but they have the right equations, so I won't cover those. Say we have a curve that corresponds to how go
by TJSomething 8y ago
Those Wikipedia pages are kind of awful for pedagogy, but they have the right equations, so I won't cover those.
Say we have a curve that corresponds to how good of a fit your model is. We want to try to find the maximum on that curve. However, calculating every point of the curve is too expensive, so we want to minimize the number of points we have to check. So, we start with a guess as to the highest point on the curve and take the first and second derivatives of the curve at that point. This gives us enough to fit a parabola that approximates the curve in the neighborhood of our initial guess point. Then, it's pretty easy to solve for the highest point on the parabola. That's our new guess. Repeat that a few times until the guesses stop shifting much. If the curve is nicely shaped (i.e. smooth everywhere, only has one maximum), the guesses will converge on the highest point.
This is a often faster than a similar method, gradient ascent, which relies upon only taking the first derivative. This would yield a line in the vicinity of our guess, and then we just move our guess a little bit, such that it goes up the line. This is pretty slow, since it can't just go straight to a guess of the top, and if you go too fast, then it'll blow right past the maximum.
The Hessian matrix is the higher dimensional equivalent to the second derivative there and the gradient is the equivalent for the first derivative. For example, if we have a two dimensional surface in 3D, then those matrices will be 2x2 and capture the curvature of the 3D paraboloid in the vicinity of the guess. As you go up in dimensionality, they're called quadric hypersurfaces.
When you're fitting a logistic regression, your hypersurface is the logarithm of the likelihood that the data you have fits the logistic curve with parameters at that point. The logarithm makes the hypersurface better behaved and makes the calculus easier. You just need the gradient and the Hessian, evaluate those at your initial guess, fit a quadric hypersurface to the guess there, pop up to the top of that hypersurface, repeat a few times, and you've got your model.
- cimmanom 8y agoThank you for a super clear explanation that was easy to grok even with no more math background than (extremely rusty) high school calculus.
- hotwire 8y agoyep, these are the kinds of posts on HN that I really appreciate, not the contrarian "well akshually" kind that usually appear.
- ghaff 8y agoUnfortunately, mathematics is one of the areas where Wikipedia is pretty awful in general. The articles seem mostly written for people who pretty much already understand the topic in question. Of course, you always have to assume some knowledge base but the stereotypical jargon-filled Wilipedia approach is particularly off-putting in this area.
- deleted 8y ago[deleted]
- Myrmornis 8y agoI understand what you’re trying to say, but Wikipedia is a fantastic resource for mathematics. “Pretty awful” is not a correct choice of words. But yes, much of it is written at beyond-undergrad-math level. And undergrad math is already advanced! And no I’m not someone with a math PhD talking down! I’m struggling through teaching myself undergrad math.
- pvg 8y agoThe only way "pretty awful" is incorrect is that it is too polite and reserved. Reams upon reams of pages are written completely at odds with Wikipedia's own style guidelines and common-sense expectations of what one might find in an encyclopedia. Unlike some famously dense mathematical texts, wikipedia maths pages don't even come with any of the benefits of brevity or focus. It's like a giant joke competition of who can describe every trivial thing in the most abstract and abstruse way except it got out of hand and the participants forgot it was supposed to be a joke. Mathworld and similar sites will help you much more with undergrad maths.
- Myrmornis 8y agoWe’re saying much the same thing, it’s just that I find your and GP’s use of the absolute “pretty awful” to be hyperbolic and something of a loss of perspective. Remember what we have here: a free, actively maintained, accurate, comprehensive and advanced corpus of expository writing on mathematics. Adjectives that are missing there are “intuitition-rich”, and “helpful for undergraduates”. I do understand if you are disappointed with it. As noted, I (undergrad level) don’t approach it with an expectation that it will be my favorite reading on a topic.
- mlevental 8y agowhat is the name of the first method (fit parabolic surface)?
- mxwsn 8y agoIn optimization it's known as Newton's method. See the third section here for an intuitive image of repeated parabola-fitting. https://ardianumam.wordpress.com/2017/09/27/newtons-method-optimization-derivation-and-how-it-works/ https://ardianumam.wordpress.com/2017/09/27/newtons-method-o... Wiki: https://en.wikipedia.org/wiki/Newton%27s_method_in_optimization https://en.wikipedia.org/wiki/Newton%27s_method_in_optimizat...
- mlevental 8y agoi feel silly for never realizing before that this was the appropriate geometric interpretation of newton's method.