3 ms·
If it helps anyone, the FiveThirtyEight article describes a scenario where people take a survey about immigration status and voting. Most legal citizens will c
by cromd 9y ago
If it helps anyone, the FiveThirtyEight article describes a scenario where people take a survey about immigration status and voting. Most legal citizens will correctly identify themselves, but some will accidentally check the wrong box and say they are an illegal immigrant. If you have a billion citizens and 10 illegal immigrants truly taking the survey, and people check the wrong box 1 in 1000 times, your "percentage of illegal immigrants who vote" statistic will be about the same as for citizens (because almost all reported illegals will be citizens). Collecting more data won't help.
It's a very good article, though in the context of deciding how many variables should be in a model of some complex phenomenon, this example is a little tougher to wrap your head around. It's not quite a predictive model, but there were some variables left out. A naive model I suppose is "this data is generated by infallible respondents", whereas a better model would incorporate that error rate. There isn't as much of a question about which pieces of information are relevant, though, like you might encounter when trying to predict future drug use from household income, race, age, number of books read as a child, number of pets, and so on.
- dopamean 9y agoIf you could find a link to the FiveThirtyEight article I'd really appreciate it. Thanks.
- foota 9y agoIt's linked in the grandparent's comment. https://fivethirtyeight.com/features/trump-noncitizen-voters/ https://fivethirtyeight.com/features/trump-noncitizen-voters...
- corey_moncure 9y agoDoesn't this make the assumption that "illegal immigrants" won't check the wrong box, intentionally or by accident?
- Double_Cast 9y agocitizens labeled illegals: 1,000,000 illegals labeled citizens: 0.01
- deleted 9y ago[deleted]
- 0xbear 9y agoIllegal immigrants have no business being anywhere near that particular checkbox.
- joshuamorton 9y agoThe other common example of the same phenomena is a test for a deadly genetic defect with a 1% false positive rate. If the incidence of the defect is .01% and you test positive, its actually more likely that you don't have the disease. (although this can be solved with bayesianism over frequentism).
- gpawl 9y agoThis is a pernicious misunderstanding of "frequentism", often found among people who studied statistics mainly by reading comics. http://web.archive.org/web/20130117080920/http://andrewgelman.com/2012/11/16808/#comment-109366 http://web.archive.org/web/20130117080920/http://andrewgelma...
- joshuamorton 9y agoThat's incredibly unnecessary. My understanding of statistics is not derived from comics (and the first time I heard that example was in a statistics course), and the link you post doesn't actually address what I stated. It addresses an actual mistake in the comic, which is a mistake that I didn't make. Here's Andrew, the author of that blog post: > Yes, I think it makes a lot of sense to criticize particular frequentist or Bayesian methods rather than to criticize freq or Bayes statisticians. Which is exactly what I did. There are times when frequentist methods are effective. I just wouldn't use them to tell me that I have a disease.