4 ms·
Pearl provides a formalism for expressing causality, however it's questionable (to me) whether that formalism corresponds to what we humans consider "causality"
by yaroslavvb 17y ago
Pearl provides a formalism for expressing causality, however it's questionable (to me) whether that formalism corresponds to what we humans consider "causality". Take 3 variables as an example, A,B,C. We observe some instances of A,B,C triples, and try to fit a probability distribution P(A,B,C) to data. If A,B,C are binary, there are 7 parameters, we can fit them directly to training data to get perfect fit, however, when we run it on separate set (validation data), it'll probably not perform very well. To improve fit on new data, you need to reduce number of parameters, so you could make some simplifying assumptions and write P(A,B,C) as P(A)P(B|A)P(C|B). In this form, there are 5 parameters to fit, and the resulting model could perform better with validation data. Alternatively you could fit P(A)P(B)P(C|A,B), 6 parameters, or P(C)(B|C)P(A|B). Select the one that performs the best on your validation data, and is hence the "best" model.
Now comes the cuasality connection -- if the best model is P(A)P(B)P(C|A,B) you read it as A->C<-B. If best model is P(A)P(B|A)P(C|B), you read it as A->B->C. (ie, A causes B, B causes C)
I think causality connection is questionable because the model that corresponds to A->B->C will have the same fit as model for A<-B<-C or A<-B->C. In fact, you could take any causal network without loops and "unshielded colliders" (connections of the form A->B<-C), pick any node as a root, and re-order the arrows to face away from the root to get a model with a different semantic causal structure, but the same mathematical structure, meaning it'll give identical fit to data.
What would be really interesting is if someone deduced causality using Pearl's approach, then verified it using a direct experiment
- cousin_it 17y agoWow. That's the kind of discussion we should be having. I didn't quite realize that Pearl's causality cannot tell A->B->C from A<-B->C. On the other hand, no other non-experimental statistical technique can do that either. Also, many problems will likely have a very constrained set of causal graphs to choose from, like smoking/tar/cancer.