7 ms·
I'm no statistician, but the whole premise seems mismatched. Why are we using a tool from regression to analyze a classification problem?
by blt 2y ago
I'm no statistician, but the whole premise seems mismatched. Why are we using a tool from regression to analyze a classification problem?
- greesil 2y agoYep. I'm also not a statistician, but linear regression that the blogger is using predicts the mean for each state, and this is being conflated with trying to predict p(color | state). The goodness of fit here would be better modeled cross entropy and not a standard deviation.
- kgwgk 2y agoFor what it's worth, the "blogger" is an statistician.
- ayhanfuat 2y agoI mean, Andrew Gelman, he is quite the statistician.
- greesil 2y agoOh he only went to MIT, pfff. And wrote a textbook.
- Xcelerate 2y agoHaha, this is why I always Google whoever the author is to articles posted on HN before commenting. More than once I've thought "this person is an idiot" only to Google their name and find out they are a famous person in that field. Then I go back and re-read their article and realize I missed something more subtle going on.
- lupire 2y agoThis seems to be a special case. The blog post was rebutted by Seth in the comments 2 weeks ago, same as Colin Percival's HN rebuttal, and Andrew didn't reply. It seems like a weird goof. Andrew was "buggin".
- Xcelerate 2y agoYeah, the more I read the post the more confused I am actually. At first glance to me this seemed like a non-paradox. So I kept wondering if I'm missing something, but based on everyone's responses, maybe I'm not?
- chipdart 2y ago> Why are we using a tool from regression to analyze a classification problem? Because classification is a regression problem. Think about it for a second. You want to put together a tool to tell which class an input belongs to. You have training data you can use to build your tool around. Your training data is already divided into sets that belong to a specific classm Your goal is to put together a model that can tell you what's the closest class your input belongs to by comparing with how close your input is to elements of the training data belonging to a specific class. What's your strategy? Well, one of the textbook strategie starts by specifying how you measure the distance between elements of your training set, and from that point you work on putting together a function that not only minimizes the distance between elements of your training set but also, when used to evaluate elements of a training set, works well in telling the type of elements of the training set that are closest to them. Then you assume the class of your input element is the same as the class of the elements of the training set that are closest to them. In the example above, the minimization step is... Yes, regression. You use regression to fit your model to your training data so that it is able how close your input element is to elements of a certain class, and then outputs how close it is to each of the classes.
- jncfhnb 2y agoClassification models break down to regression problems under the hood, but regression metrics are not good tools to evaluate the efficacy of classification models.
- chipdart 2y ago> Classification models break down to regression problems under the hood, but regression metrics are not good tools to evaluate the efficacy of classification models. You're simply wrong. Regression is a tried and true classification technique. Posting personal and baseless assertions don't refute that. I mean, pick up pretty much any textbook on supervised and unsupervised learning. You always end up with an approach which boils down to having training data, put together a trial function, apply a minimizer to fit trial functions to training data, and evaluate the resulting model by running trial data through it. Fitting trial functions to training data has a name: regression. Minimum squares has a very precise interpretation both in linear models and in probability. There is no way around it.
- vcdimension 2y agoI am a statistician, and you're right, for this kind of thing we would normally use a binary response model such as a logit or probit model that constrains the response variable to be between 0 & 1. However in this case it doesn't matter since there's only one independent variable (state), and it's binary so there's only 2 different predictions the model could make (which will be the correct probabilities of 0.45 & 0.55, even with a linear model). The normal R^2 formula can't be applied to a logit/probit model; instead you use an alternative such as McFadden's or Cox & Snell pseudo R-squared. I'd be interested to see what value they take for this example. Linear models are sometimes used even in models with many independent variables since it can be shown that the coefficients in a linear model are unbiased estimators for the average partial effects of any non-linear binary response model.