3 ms·
> First off, why are assumptions independent? Because I've defined them that way. I mean them to be independent choices you could make when designing your mode
by tbabb 7y ago
> First off, why are assumptions independent?
Because I've defined them that way. I mean them to be independent choices you could make when designing your model that could be varied to fit the data. If two aspects of the model are not independent; i.e. they are covariant in some way, then there is some common parameter that explains them both, and that parameter is the one that should be seen as an input to the model.
> We can roll a die 100 times, 1000 times, 10K times [...] That's what we mean (if we're frequentists)
We're not frequentists.
You can't "re-roll" the 2016 election 10K times, either. There was only one, and there was only one way it could come out; we just didn't know enough to say what it would be before it happened. All the particles in all the voters were obeying the laws of physics at every moment; never was there any freedom for a different outcome. Nonetheless, even though there was/is only one "ground truth" that could ever be, we assigned probabilities to each possible outcome, given our incomplete knowledge.
This is a pretty standard application of probability. State estimators (e.g. the Kalman filter) are doing the same thing— you have some noisy readings of reality, and you use Bayesian logic on some assumed probability distributions to pick the estimate from the space of possible "ground truths" that has the highest probability of being the right one.
Concretely: I'm measuring roll rate, local acceleration, compass heading, barometric pressure, and GPS, all with significant error, and I want to know where my quadcopter is most likely to be at the current moment. There is only one true answer to that question, the quadcopter is in one place, not 10,000 places (or 10,000 flights), there is a single ground truth. But Bayes will give me a probability, given my readings, that any given estimate is the true ground truth (and some math will help me solve for the highest one).
In this case, instead of assigning probabilities to possible election outcomes or system state "ground truths", the "configuration space" is models of reality. But all we've changed is the domain of our probability distribution; the math doesn't care what kind of thing our "ground truth" represents. And it doesn't matter if reality contains only one "ground truth" or many; the fact is that we are choosing between many options (and we are ranking them by likelihood).
- vladf 7y agoRe independence: identifying whether or not these physical assumptions covary is not that easy. That's my point: assumptions a, b, c, d could easily have some mutual incompatibility that makes them non-independent. It's an active area of research. Re probability, I'm glad you committed to the Bayesian interpretation. Bayes gives you a _degree of belief_, based on your priors. It's quite fortuitous that you mention the 2016 election. As you say, there's only one instance here. Which is why the prior matters a lot. We can incorporate (partial) evidence from past elections, but it's going to be very sensitive to the priors that we place, since the net amount of evidence we're working with is very small. As we found out in 2016, that means these beliefs aren't worth much in such low data scenarios, since the prior has a large impact! https://projects.fivethirtyeight.com/2016-election-forecast/ https://projects.fivethirtyeight.com/2016-election-forecast/ This brings me to my original point: > In the limit of evidence, this prior matters less, but constants matter here! How much evidence do we need before we can be confident the "belief levels" we're throwing around aren't that subjective anymore? We don't really have a good sense for what the structure over this space of "assumptions of physical models" is, so we can't really answer this question.
- tbabb 7y ago> identifying whether or not these physical assumptions covary is not that easy But still tractable, I'd say. My core claim is that counting independently-variable assumptions will be a highly performant way to select between theories which agree with the data. Or put another way, it's the best Fermi approximation calculation for measuring "how good is your theory". How you do that for any given theory, while important to do correctly, is an implementation detail which I think is secondary to the discussion of whether doing it at all is a good idea. :) (seems like we might agree that it is?) > We can incorporate (partial) evidence from past elections, but it's going to be very sensitive to the priors To the extent that election forecasts are unreliable, I think that's because they are forced to involve a lot of assumptions (e.g. similarity to past elections) that turn out not to correspond well to reality. Models which make fewer such assumptions will do likely do better! (and IMO fivethirtyeight's forecasts did the best job of this out of any; most of the rest put Hillary at around 97%). Unfortunately with elections, there is a comparatively high lower bound on the number of assumptions we must make, thanks to the complexity of their dynamics and sparsity of data/knowledge we have about each. I think this is much less the case with physics, where we are varying comparatively small physical assumptions to explain mountains of data. But in either case, I contend that the most performant models will make fewer (unmeasured) independent assumptions. > How much evidence do we need before we can be confident the "belief levels" we're throwing around aren't that subjective anymore? The point I'm trying to make is bigger-picture than the above level of detail: Counting independent assumptions, in the limit, matters more than the specific constants of each assumption (assuming they're not low/zero), precisely because it's so hard to come up with "accurate" numbers for each. That is to say, the probability is not sensitive to those belief levels: We could choose widely varying distributions for the probabilities of our assumptions, including choosing probabilities very close to 1, and it will hardly ever matter to the total probability as much as the absolute number of independent assumptions we make.