5 ms·
You look across unrelated events that you scored similarly. In aggregate, across all things you predict as having an 8% chance to happen, 8% should actually hap
by dadrian 3y ago
You look across unrelated events that you scored similarly. In aggregate, across all things you predict as having an 8% chance to happen, 8% should actually happen.
- jemmyw 3y agoThat sounds reasonable. The article doesn't dig into it. Does this group actually make enough predictions to have this kind of validation?
- draaglom 3y agoYes, serious forecasters predict on hundreds of questions & the relevant websites prominently feature "binary calibration" plots. This is mine on Metaculus (mentioned in the article): https://imgur.com/a/d3s67xk https://imgur.com/a/d3s67xk There's still a bit of noise at 205 predictions but you can see a pattern emerging!
- anonymous-panda 3y agoBut if you can pick what you’re predicting on and pick the percentage you assign to an event, I fail to see how that gives you a calibration.
- tunesmith 3y agoBecause you make lots of predictions, and smooth out the numbers. If the events you give 8% probability actually happen 35% of the time, you aren't as calibrated as you are if they happen 10% or 5% of the time. The calibration numbers measure how much you are off by.
- anonymous-panda 3y agoThis makes no sense. If I have 1000 predictions and vary my estimates from 0-10%, that’s only 100 samples assuming you round to the nearest whole percent. And there’s no correlation between any of those samples. For example, I could say the probability of a lightning storm tomorrow is 1% and the probability of a war with China in the next year is 1% - calibrating amount those 1% events is clearly non sensical. You could try to calibrate among similar events but then you have no way of estimating that a 10% war between US and China vs a 1% of war between Russia and China. Basically this is an exercise of garbage in and garbage out.
- omeze 3y agoI think youre kind of right, its not really meaningful to know “im right about things that happen 5% of the time!”. But if a lot of things you bet on structurally happen to be low probability and of the same type (eg you bet on one type of catastrophic event like pandemic, crop disease, or war etc) then its useful
- anonymous-panda 3y agoOr you’ve just learned how to optimize your score on the game you’re playing but the game isn’t actually about predicting the future.
- tunesmith 3y agoProper scoring rules like Brier mean that any attempt to optimize your score by definition mean that you are improving your calibration.
- hackerlight 3y agoThe article does dig into it. They have Brier scores.
- jellyberg 3y agoSee https://samotsvety.org/track-record/ https://samotsvety.org/track-record/
- dotancohen 3y agoThat sounds a bit like the gambler's fallacy to me.
- avereveard 3y agoDoesn't seem right, if a state capital city base chance for a nuclear strike is 1% and you have 195 capitals or so 2 of them would be hit by now.
- TOMDM 3y agoYou're correct it mostly only works for a spread of predictions that don't have such a large influence on each other. If one capital gets nuked the odds of another one suffering the same fate skyrocket.
- dmurray 3y agoYeah, it's just about possible to imagine Kyiv or Tehran gets nuked in a "limited engagement" that never escalates further. But if, say, Lisbon gets nuked then the whole world is at war.
- cudder 3y agoMaybe your base rate is way off then?
- danmaz74 3y agoNuclear strikes on N different capital cities aren't statistically independent events, and that kind of computation only makes sense for statistically independent events.
- michaelt 3y agoThat might give you some confidence. But if you're looking at your performance on unrelated predictions, it won't necessarily detect problems with a given prediction. I predict there's an 8% chance of one drive in my four-hard-drive array failing in the next year. I predict there's an 8% chance the next UK election ends without any party holding a clear majority, leading to a labour-lib dem coalition government. For the first prediction to be accurate, I just need to read the backblaze hard drive stats and multiply. For the second prediction to be accurate, though? That depends on a lot of factors that are a lot harder to know. For example, would the lib dems be likely to enter a coalition, given how the last one went for them in 2015?
- yorwba 3y agoSure, you cannot expect to reliably detect problems with a given prediction, since it could always be right or wrong due to chance. But also, what would you do if you had the ability to detect problems with a prediction? If you just want to use a different prediction method that fixes the problem and makes better predictions, looking at aggregate performance over many unrelated predictions does help. Compare how well different methods do, go with the best.
- harperlee 3y agoBut you could come up with a probability prediction on whether a random prediction from the predictor is accurate.