4 ms·
I'm a big fan of Nate Silver, but I think this idea that he "predicted N/50 states correctly" in 2008 and 2012 is mostly a myth. (1) The website doesn't "predi
by keithwinstein 3y ago
I'm a big fan of Nate Silver, but I think this idea that he "predicted N/50 states correctly" in 2008 and 2012 is mostly a myth.
(1) The website doesn't "predict" a state when it reports that the probability of (e.g.) an Obama victory in Florida is 50.3%, which I believe is what 538 assessed in 2012. If you think that a state's probability distribution is 50/50 between Candidate X and Y, and the outcome is Candidate X, obviously you don't get any points for a good prediction. You didn't make a prediction. And if you say it's 50.3% probability, it's essentially the same situation. You didn't make (much of) a prediction. Which is okay!
(2) In terms of Brier Scores or entropy scores, which compare the probability with the outcome and give more "points" the stronger the prediction was made (vs. what actually happened), FiveThirtyEight's predictions are generally just middle-of-the-pack compared with other websites and mechanisms of aggregating these polls. (See, e.g., https://web.math.princeton.edu/~sswang/Wang_Origins_of_Presidential_poll_aggregation_withfigures.pdf https://web.math.princeton.edu/~sswang/Wang_Origins_of_Presi...) There's no strong evidence that 538's predictions are much better or worse than other reasonable ways to weight and aggregate the recent polls that anybody else might come up with.
(3) And... almost everything that tries to quantify the accuracy of the predictions is talking about the state of the 538 website ON THE MORNING OF THE ELECTION. But that's probably the least useful day to be looking at FiveThirtyEight. By then, you can just wait a few hours and get the actual results. :-) The predictions that are really valuable are the ones that are made some time in advance, e.g. one or two months.
But (a) it's hard to directly measure the "accuracy" of a prediction like that, because... things change and we can't rerun the universe, and (b) it's not like the model is somehow time-invariant to be able to measure the accuracy indirectly by proxy; there are various "time until election" factors included in these model itself, the dynamics of the election itself probably have a "time until election" dependency, AND the website operators, including 538, are generally manually tweaking the model parameters or behavior in ways they don't always disclose as the election evolves.
So basically... it's hard to say. I think 538 has done great work, and in general it's incredibly useful to have a consistent world model that you can rationally update to discuss the effect of recent news. The FiveThirtyEight political models are as good as any for that! But I don't think there's much strong evidence they're better.
- saghm 3y agoThat's fair! I probably should have realized that, given that I already knew that his 70-30 prediction of Clinton winning in 2016 didn't mean he "predicting wrong"; I'm guessing I was just less aware of how his models worked back in 2012 than I was by 2016, so I didn't really consider what it meant to "predict the states".
- yowzadave 3y agoI think many people subconsciously conflate a "70% chance of winning" with a "70% of the vote" prediction, when the reality is much closer--a 70% chance of winning could very easily go the other way, given some systemic polling error and/or late-breaking news!
- saghm 3y agoYeah, it's definitely tricky to communicate. The chances of _either_ candidate getting 70% of the vote in any recent presidential election would be incredibly small; even 55% would be a huge anomaly. Most people wouldn't feel very enlightened looking at a probability distribution of potential vote percentages for each candidate on a CNN graphic though, especially when they'd basically just look like giant spikes around 48-52%.
- dragonwriter 3y ago> But it’s hard to directly measure the “accuracy” of a prediction like that, because… things change and we can’t rerun the universe Its hard to measure the accuracy of a prediction like that, but with a big collection of predictions (and since every cycle Silver predicts a lot of individual elections – not just each of the state elections for presidential electors, including both the statewide electors and the by-district electors where those exist, but also house, senate, gubernatorial, and some other races – each election cycle, this exists) you can bucket them in similar probability buckets and say, for instance, and see how often things Silver predicted in each probability bucket came true vs. the probability range of the bucket, and get a good idea of the accuracy of the predictions.