5 ms·
I'm sorry, that seems like the most backwards explanation of logistic regression I've ever seen. It uses the sigmoid function because it's logistic regression.
by oddthink 10y ago
I'm sorry, that seems like the most backwards explanation of logistic regression I've ever seen. It uses the sigmoid function because it's logistic regression. The statistical model is a Bernoulli distribution. Write down the likelihood, maximize it, and you get the sigmoid.
- llasram 10y agoI think this is a good example of the problems with autodidactism. For all its flaws, structured education is what makes common knowledge common. When you study on your own, you don't know what you don't know, and there's no one there to point out your obvious-with-the-right-knowledge lapses.
- brudgers 10y agoTo me it's an example of what Hacker News offers people interested in learning [the human, not machine type]. This very article shows how it is a place where there are no long term consequences equivalent to grades: just a wee bit of standard internet asshattery and those complaints about the youth of today that have been in the Western cannon since Socrates was drinking Hemlock.
- krosaen 10y agoYes, it uses a sigmoid function because it's logistic regression and therefore you are using the inverse of the logistic function, the sigmoid, as the notebook explains. But I think it's worth running through that and exploring why it's useful to use a logistic function in the first place (maps linear combo to (-1, 1) range). It was useful to me anyways, perhaps obvious to others. Thanks for pointing out the relation to the Bernoulli distribution.
- ryanmonroe 10y agoThe logistic function maps to (0,1). The hyperbolic tangent tanh(x)=2*logistic(2x)-1 maps to (-1,1).
- brudgers 10y agoThe reverse nature might be due to approaching the problem from practical application of machine learning technology rather than the higher level abstracts of than from mathematics/statistics. Which is to say that the starting point is that of someone figuring out what is going on behind the scenes of a working software package rather than building up from first principles. Curious if there is a link to a better explanation from a similar starting point.
- gbrown 10y agoThe real explanation comes from the theory of generalized linear models: https://en.wikipedia.org/wiki/Generalized_linear_model https://en.wikipedia.org/wiki/Generalized_linear_model
- brudgers 10y agoMy apologies for not being clear, I was wondering if there was something closer toward the Randall Monroe [edit] "thing explainer" end of the spectrum. The reason I was wondering is because machine learning is becoming something that people incorporate into a project as a library. The analogy I would draw is to something like TCP/IP where there are abstractions over routing and congestion control and where routing and congestion control are abstractions over the mathematics of scheduling and graph theory. Maybe a better analogy might be cryptography where practical implications of entropy pool design are relatively esoteric despite the vastly less accessible mathematical nature of reliable one way encryption algorithms.
- hiddencost 10y agoRandall Monroe?
- brudgers 10y agoThe training set for "randall" sigmoid in my brain becomes apparent.
- ogrisel 10y agoThe statistical model is a Bernoulli distribution where the expected value p is parameterized by the logistic sigmoid function applied to a linear model of the input variables. The choice of the use of the sigmoid function does not stem from applying MLE to the model: instead it is an priori and arbitrary modeling decision that has been done even before starting to think about the estimation of the parameters of the model. The intuitive justification with respect to the choice of this link function given in this blog post seems quite standard to me. This text-book chapter gives a similar intuition: http://www.stat.cmu.edu/~cshalizi/uADA/12/lectures/ch12.pdf http://www.stat.cmu.edu/~cshalizi/uADA/12/lectures/ch12.pdf (section 12.2). Also in practice we don't use pure MLE but l2-penalized (or l1-penalized) MLE to fit a logistic regression model as otherwise it might be very sensitive to noisy data if the number of features very large (compared to the number of training samples).
- oddthink 10y agoYep, you're right. I was playing too fast and loose there. The Bernoulli likelihood definitely suggests the odds as a natural space, but it's not required.
- krosaen 10y agoThanks for pointing out the similarity in intuitive justification in Shalizi's text, I have that book on my to-read list already, fun to see I wasn't far off.
- fiatmoney 10y agoYou can actually do it in either direction, going from the link function and reasoning back to the distribution it implies, or starting at the distribution and seeing what maximizes the likelihood. The "logistic regression" has been invented / reinvented independently with different derivations 10 times or so.
- stdbrouw 10y agoIf you look at the history of statistics, you'll find plenty of practically motivated ad-hoc problem solving like "hmm, how do we get R^n to map onto [0,1]" Log odds were used long before logistic regression and GLMs came along, and it certainly doesn't sound backwards to me to try and explain how people came to use a technique that initially looks very arbitrary (why not model probabilities directly?)