4 ms·
Interesting article and I find it current for some problems I'm working on at the moment. I would add a few challanges. The example is a bit a of a strawman -
by graeham 10y ago
Interesting article and I find it current for some problems I'm working on at the moment.
I would add a few challanges. The example is a bit a of a strawman - a log(x) function has unique properties that make the Xmax-Xmin vs R^2 work like that. In real data, rarely does a single-variable 'true model' fit as well as the example either.
Context is needed as well - depending on the use of the model, a linear or quadratic fit may be sufficient even for what is clearly a log dataset. The real failing on only for small values of x, maybe 5% of the range of total values. For this case, a bilinear model could fit quite well for the lower 5%, then the existing model for the upper 95%. It depends on the application. I like this phrase:
"When deciding whether a model is useful, a high R2 can be undesirable and a low R2 can be desirable."
Too often statistics are dominated by 'cutoff' values that people apply blindly to all situations.
What do you think of robust regression methods, where obvious outliers are down-weighted?
- yummyfajitas 10y agoI'm not the author, but I'm a huge fan of robust regression. I make between $500-2000/month off a trading strategy based on such a method. (The method is basically Bayesian linear regression, but using an error model that has a heavier tail than a gaussian.) But a really important thing when using such methods is the lucas critique. When you need to use robust regression you are definitively in a space where all the simple and generic stuff (e.g. linear regression) doesn't work. So at this point it's important to validate the underlying assumptions behind the robust regression scheme. E.g., in my trading strategy, I've gone to great lengths to make sure the tail behavior I'm modelling is an overestimate of reality.
- xoranth 10y agoCould you share more details about the robust regression you are using? All resources I could find online on robust regression would either point to Laplace distributed residuals, or some capped loss function.