4 ms·
> https://sci-hub.3800808.com/10.1038/302687a0 https://sci-hub.3800808.com/10.1038/302687a0 I'm having trouble figuring out a non-absurd interpretation of that
by _Nat_ 4y ago
> https://sci-hub.3800808.com/10.1038/302687a0 https://sci-hub.3800808.com/10.1038/302687a0
I'm having trouble figuring out a non-absurd interpretation of that paper.
For example, their Equation-7:
> p(h←e, e) = p((h←e)e, e) = p(he, e) = p(h, e)
, which looks like they're saying that, when there's evidence "e", the probability that a hypothesis "h" is true is equal to the probability that "e" proved it.
For example, say we consider the hypothesis, "h", that there aren't 10-armed spider-monkeys that like jazz-music currently on Earth. Then based on that hypothesis, we make the prediction that we won't see a 10-armed spider-monkey listening to jazz-music in the next room. Then, let's say we check that room, and there's not such a 10-armed spider-monkey listening to jazz-music, such that there's "e".
Did our test prove the hypothesis? Wouldn't seem like it.. I mean, even if there were 10-armed spider-monkeys, it'd seem like they could just be in places other than in the other-room. So, "p(h←e, e)" would seem pretty close to zero.
However, it still seems like the hypothesis that such spider-monkeys don't currently exist on Earth would seem fairly probable. So, "p(h, e)" would seem fairly close to one.
So, "p(h←e, e) = p(h, e)" wouldn't seem to hold, even approximately.
That said, the author didn't clearly specify exactly what they meant, so maybe they meant something else? But I'm not seeing an obvious, non-absurd interpretation of their claims.
- _Nat_ 4y agoFollow-up: I think I found some relatively unambiguous claims in the paper to help get a handle on the logic. After "Theorem 2", there's a "Proof" with multiple lines directly equated. Among those lines were > = 1 - p(h,eb) - p(e,b) + p(he,b) > = p(h←e,b) - p(h,eb) , which can be equated to and reduced to find 1 = p(h←e,b) + p(e,b) - p(he,b) , and then if we take the condition of "b" as assumed for brevity, 1 = p(h←e) + p(e) - p(he) , then it appears that the conditions of "h←e" and "e" cover all possibilities, plus an excess overlap of "he". So, "h←e" refers to NOT(e) plus AND(h,e). So, "h←e" equals OR(NOT(e), AND(h,e)). So, the evidence "e" implies the hypothesis "h" when both are true, plus also when evidence "e" is false. --- So, "Theorem 1" claims p(h←e, e) < p(h←e) , which we can now parse given the above to OR(NOT(e), AND(h,e)) when e < OR(NOT(e), AND(h,e)) , and we can reduce the left-hand side to find h when e < OR(NOT(e), AND(h,e)) h when e < NOT(e) + AND(h,e) h when e < NOT(e) + (h when e) * e 0 < NOT(e) + (h when e) * e - (h when e) 0 < NOT(e) + (h when e) * (e - 1) 0 < (1 - e) + (h when e) * (e - 1) e - 1 < (h when e) * (e - 1) 1 - e > (h when e) * (1 - e) 1 > h when e , or to write that last line out, p(h | e) < 1 , which matches out with the condition that they attached to "Theorem 1", which requires that p(h|e)!=1. But to work that out with the sides keeping their values, h when e < NOT(e) + (h when e) * e h when e < NOT(e) + (h when e) * (1-NOT(e)) h when e < (h when e) + NOT(e) - (h when e) * NOT(e) h when e < (h when e) + NOT(e) * (1- (h when e)) h when e < (h when e) + NOT(e) * (NOT(h) when e) , which appears to be the last line of their "Theorem 2". So.. I guess that explains the definitions that they were using. --- Anyway, what seems odd to me about that is that "Theorem 1" seems like it's meant to be surprising -- like it's meant to show that finding evidence reduces the meaningfulness of the evidence itself, or something? However, some things seem off. For example, the expression of "h←e" seems weird to me; it'd seem more sensible for it to be like this: OR(AND(NOT(e), NOT(h)), AND(h,e)) when e < OR(AND(NOT(e), NOT(h)), AND(h,e)) h when e < OR(AND(NOT(e), NOT(h)), AND(h,e)) h when e < (!h when !e) * !e + (h when e) * e 0 < (!h when !e) * !e + (h when e) * (e - 1) 0 < (!h when !e) * (1 - e) - (h when e) * (1 - e) 0 < (!h when !e) - (h when e) (h when e) < (!h when !e) , where the inequity isn't obviously of particular interest. Because the second thing that seems off is the notion that this matters -- that the evidence, "e", should be a concern for not just figuring out the probabilities in the model, but also retro-actively adjusting the meta-model, or something? In short, after tracing their math and such, it's unclear what point they might be trying to make, as this doesn't seem surprising or unexpected.
- MatteoFrigo 4y agoThe way I read it, the comma "," means "given", so what they write p(a, b) would be written as the conditional probability p(a|b) in contemporary CS/math literature. h <- e means (h OR (NOT E)), which is the usual meaning of implication (either the consequent is true or the antecedent is false) So p(h <- e GIVEN e) = p((h OR (NOT e)) GIVEN e) = p(h GIVEN e) since NOT e is false given that e is true.