4 ms·
We wrote a couple docs about regression. I'd be curious to get folks' feedback. Guide itself: http://docs.statwing.com/user-friendly-guide-to-regression/ http:
by glaugh 11y ago
We wrote a couple docs about regression. I'd be curious to get folks' feedback.
Guide itself: http://docs.statwing.com/user-friendly-guide-to-regression/ http://docs.statwing.com/user-friendly-guide-to-regression/
Perhaps more interestingly to a lot of folks in this crowd, a guide to interpreting residuals: http://docs.statwing.com/interpreting-residual-plots-to-improve-your-regression/ http://docs.statwing.com/interpreting-residual-plots-to-impr...
One valid critique is that the approach we describe is not super rigorous, there's more "add a variable and see if it sticks" and "explore a bunch of bivariate relationships to decide what to include" then you'd want to do if you were publishing a paper on your findings. To our minds, though, this approach is more realistic for the use cases we're more concerned about, like analyzing an ad hoc survey. You can always consider your results to be exploratory and then validate later. (Note that these docs occasionally refer to Statwing, our product, but really could be used for any tool).
Criticism is welcome.
- fats_tromino 11y agoJust as two quick comments, confidence intervals != prediction intervals. One gives a range around a true mean value, the other gives a range around where the variable may actually fall (prediction intervals are bigger). You may also want to mention adjusted R^2 as a measure of quality of the model. I've never heard of AICR before, the standard metrics for quality of the model are AIC/BIC (and sometimes Mallus Cp). Edit: fixed sloppy wording about confidence interval.
- stdbrouw 11y agoAIC depends on residual variation just like R^2, so I guess some people might be inclined to call it AIC_R to make that clear. Also, it's Mallows' C_p :-)
- fats_tromino 11y agoThanks for the clarification about Mallow's C_p, I always get the name messed up for some reason.
- fnbr 11y agoHey, just FYI, your posts are getting killed for some weird reason. See: http://imgur.com/FmBkqgo http://imgur.com/FmBkqgo
- fats_tromino 11y agoThank you, I'm confused as to why this is happening.
- glaugh 11y agoThe prediction interval thing was sloppy, good catch. The r^2 and AICR comments reflect the fact that these docs were built around Statwing's regression capabilities, which default to m-estimation instead of OLS. There's no adjusted r^2 for that, which I agree is a better measure when available, and it uses AICR, where the R stands for "robust". But still a good catch, since I didn't caveat ahead of time that we were only talking about robust methods (and we don't note it in the docs). Much appreciated.
- hooloovoo_zoo 11y agoOLS is an M-estimator.
- hackaflocka 11y agoDo you have a doc on whether the data are supposed to be Normally Distributed, and if they are then how to measure said Normality? I'd really appreciate such a doc if you do.
- glaugh 11y agoSorry, we don't. My understanding, though, is that the typical statistical measures of normality (e.g., Kolmogorov–Smirnov) aren't really that effective, and the best way to assess normality is a visual inspection of a Q-Q plot. I don't have a specific source for this assertion, it's my memory from having researched this stuff a few years ago.