6 ms·
What do you mean by coordinate transformation? MLE is invariant under parameter transformations because it's just the argmax of the likelihood.
by Akababa 7y ago
What do you mean by coordinate transformation? MLE is invariant under parameter transformations because it's just the argmax of the likelihood.
- jules 7y agoIndeed, it is the argmax of the likelihood, but the likelihood is not invariant under coordinate transformations. The quantity p(x)dx is invariant, not p(x). By picking a suitable coordinate transformation you can put the MLE on any value where the likelihood is not zero.
- deleted 7y ago[deleted]
- knzhou 7y agoAdding to the other comments, you still have prior-dependence on a more subtle level, because it depends on what hypotheses are allowed. Here's an extreme example. Consider flipping an apparently fair coin and getting "THHT". The hypothesis that the coin is fair gives this result with likelihood 1/16. The hypothesis that a worldwide government conspiracy has been formed with the sole purpose of ensuring this result... has a likelihood of 1. But nobody would ever declare this the MLE, because "government conspiracy" isn't one of the allowed options. But it isn't precisely because it's unlikely, i.e. because of your prior. Of course this is an extreme example, but there are more innocuous prior-based assumptions baked in too.
- eanzenberg 7y agoAgain: Priors can and are used to mislead. Both methods can and are used to mislead. Just moving to Bayes doesn't assume the finding is free of bias all of a sudden.
- c2471 7y agoIt doesn't. But the workflow of Bayes forces you be explicit. If you try and cook the books, it will be shown for the world to see. Can you provide a paper that quoted a p value for a regression and also validated all the asymptotic conditions are close to being true in order for that p value to be even somewhat reliable?
- closed 7y agoWait, in frequentist statistics getting, say, a p-value of 1 is not a bad thing--unless you erroneously assume that value is evidence for your null hypothesis. Consider that if your data generating process really is a fair coin, then the conspiracy outcome you mention only occurs 1 our of 16 times, so 15 out of 16 times you observe a likelihood of 0. 15 out of 16 times your reject the conspiracy case. There is also a tricky component here, because the notion of sample size is not clearly defined (can we generate multiple 4-tuples of flips, and consider each one a sample? Is your example really just a funky way of discussing type II power?)
- knzhou 7y ago> Wait, in frequentist statistics getting, say, a p-value of 1 is not a bad thing--unless you erroneously assume that value is evidence for your null hypothesis. That's exactly what I'm saying. Suppose you get HHTHT. Then you run the following statistical test: Hypothesis: a government conspiracy has been hatched to make you get HHTHT. Null hypothesis: this is not the case. The p-value is 1/32, so the null hypothesis is rejected. This is bad reasoning for two reasons: first the alternative hypothesis is incredibly unlikely, and second the choice of alternative hypothesis has been rigged after seeing the data. These are exactly the two reasons so many social science studies running on frequentist stats have done terribly, and why we would benefit from Bayesian stats which force you to make these issues explicit.
- eanzenberg 7y agoIt's strawman to always posit frequentists as unthinking blobs of meat who don't consider the credibility of the alternate hypothesis. In fact, many experimental scientists, physicists, biologists etc. made discoveries using frequentists techniques that didn't rely on boogyman notions of "want to bet the sun just burned out because you're in a closet" nonsense.
- knzhou 7y agoI'm a physicist that uses frequentist statistics, and it works fine. However, it can't be denied that some fields misuse it, though precisely the failure modes I pointed out.
- lottin 7y agoNot sure I follow? The hypothesis that the result you see is the result a worldwide government conspiracy is 100% supported by every result that you see. Because it is 100% consistent with the data, a statistical analysis will tell you exactly that--that it is 100% consistent with the data.
- kgwgk 7y agoMLE is not invariant under parameter transformations because it's just the argmax of the likelihood! Take for example x~normal and exp(x)~lognormal. The maximum of the distribution is at mu for the former and at exp(mu-sigma^2) for the latter, instead of exp(mu).