4 ms·
It's not that difficult an idea, even if the trend of comment voting here indicates a disturbing ignorance of how statistics can be manipulated. What would be
by WilliamLP 17y ago
It's not that difficult an idea, even if the trend of comment voting here indicates a disturbing ignorance of how statistics can be manipulated.
What would be mathematically sound would be to agree on a standard set of tests, perhaps including Benford's law, digit distribution, two digit combinations shown to be popular when humans pick numbers at random. And to do so _BEFORE THE DATA HAVE BEEN AVAILABLE_.
How would I trick people into believing a false conclusion? One way would be to find 30 or so measures of what could suggest suspicious patterns. Then I'd test the data with them, and perhaps there would be 2 which have a 5% chance of occurring purely by chance. Then I say "Aha! Look at this! The chance of these two things happening are .05 * 0.05 or practically impossible!"
Apparently I could fool some pretty smart people this way, since they obviously aren't immune to basic statistical fallacy.
As another poster said, none of this suggests that the data weren't manipulated, but just that the stats in the article, as given, are purely bullshit.
- kvh 17y agoFor instance, the same authors, in a previous paper analyzing nigerian election results, state the following "lab experiments indicate that individuals tend to favor small numbers, even when subjects have incentives to properly randomize. Second, individuals underestimate the likelihood of digit repetition in sequences of random integers, so we should observe relatively fewer instances of repeated numbers in manipulated vote tallies." there is no mention of either of these statistics in the Iran analysis.
- itistoday 17y agoYou raise an interesting point, but I'm not sure I buy this argument. What is the chance that your scenario is realistic? In other words, what is the chance that given a set of legitimate vote counts, and a set of 30 known legitimate authenticity tests, that 2 of those tests would report a value of 5%? It could turn out that if the set of 30 tests are all "good tests", then the chance of your being able to find 2 tests that report low numbers like that approaches 0%, in which case you would simply be wrong. I am not a statistician, and statistics is notoriously difficult for humans to deal with [1], so until I find an answer to the question above I'm going to maintain a neutral position in this argument. Unless you happen to have a doctorate in statistics, in which case I will take your word for it. ;-) [1] TED talk: http://www.ted.com/talks/peter_donnelly_shows_how_stats_fool_juries.html http://www.ted.com/talks/peter_donnelly_shows_how_stats_fool...
- jibiki 17y agoIt's easier to compute the probability that less than 2 tests report a value of 5%. This is: (30 choose 0)*(.95^30) + (30 choose 1)*(.95^29)*(.05^1) = 0.553542075 So there's about a 50-50 chance (obviously I'm assuming that the tests are independent of each other.) The relevant Wikipedia article: http://en.wikipedia.org/wiki/Binomial_distribution http://en.wikipedia.org/wiki/Binomial_distribution
- itistoday 17y ago"less than 2"? Assuming you're doing the math correctly, it means that there's a 50-50 chance that "0 or 1" of the tests "report a value of 5%" (and does that mean less than or equal to 5%?). Which, again, doesn't say very much... The question was: What is the chance of being able to find exactly 2 tests?
- jibiki 17y ago> The question was: What is the chance of being able to find exactly 2 tests Sure thing. That's: (30 choose 2) * (.95^28) * (.05^2) = 0.258636738 I thought you meant "at least two tests", for obvious reasons. (If 4 tests showed vote rigging, the researchers would still report vote rigging...) > (and does that mean less than or equal to 5%?) Yes.
- itistoday 17y agoThanks, looks like I'll have to brush up on my stats to analyze this further... But even without direct knowledge of how this works, shouldn't the "goodness" of the tests play a role in the numbers you're using? It doesn't seem to... > I thought you meant "at least two tests", for obvious reasons. Yes, sorry, I did, I guess you meant to say I should subtract 0.553542075 from 1 to get my answer. (just saw earl's post above...)
- earl 17y agoYou are, frankly, speaking nonsense. As you pointed out, the tests will assuredly not be independent of each other. Second, what you have calculated is... I don't know. You are calculating the cumulative distribution function for a binomial distribution with n=30, p=0.95. How on earth does that relate to the previous post? Are you somehow confusing a p-value -- which is a statement about the minimal alpha for a rejection region such that given a set of data X and a test, we would reject our null hypothesis H_0 -- and a probability? Because they have almost nothing to do with each other.