3 ms·
This comment is surprising to me. Most data scientists use results and measures of statistical significance provided by the program they're using, which account
by kxxsc 7y ago
This comment is surprising to me. Most data scientists use results and measures of statistical significance provided by the program they're using, which accounts for the distribution used. Do you have examples where either data scientists are not presenting aggregate statistics or where someone is using the wrong kinds of p-values?
To your specific examples - the coefficient of a linear regression is distributed normally (?). Similarly, we know the expected distribution of most maximum likelihood estimators (logistic regression, etc.), and programs will give you the right p-value.
Of course, omitted variable bias is still a problem and it is possible to mis-specify your model. However, I think most data scientists are presenting aggregate statistics (means, regression coefficients) like you said, and that we have a pretty good handle on the underlying distributions.
- throwlaplace 7y ago>which accounts for the distribution used what does this mean? are you trying to say the code uses the empirical distribution as a proxy for the true distribution?
- wbl 7y agoMispecification error is hard to deal with.
- eanzenberg 7y agoFor example, there were times where features were excluded because their "p-value was too large to be significant", regardless that the underlying distribution was not Gaussian. p-value from a t-test, like in lots of regression software, requires Gaussian distributions.
- QuesnayJr 7y agoIf the data set is large, then the coefficient is approximately Gaussian, because of the CLT. This is one appropriate setting to use the CLT (unless you think you are in a setting where the CLT doesn't hold, such as infinite variance).
- eanzenberg 7y agoFor example, lets say height is a feature in your model. No matter how big the size of data, it will never be Gaussian, it is bi-modal. So the t-test in regression won’t be valid. Most ml is done on raw data.
- QuesnayJr 7y agoIf the t test is on a regression coefficient, then the sampling distribution is approximately Gaussian (for big enough data). It doesn't matter how many modes that the original data feature has. This is standard asymptotics in hypothesis testing.
- eanzenberg 7y agoNo, a t-test makes no assumptions that the underlying data is Gaussian. Again, most ml is done on raw data. If the raw feature is bimodal then the raw data is bimodal.
- QuesnayJr 7y agoI can't figure out if you are agreeing or disagreeing with me. If you do non-penalized regression on raw data, then the t statstic will be approximately Gaussian, even if the raw data is bimodal. This follows from the CLT.
- eanzenberg 7y agoWhat is the standard deviation of a bimodal distribution with modes at 0 and inf? Is that a meaningful stat?