4 ms·
I mostly agree with Pearl. However, the account of traditional statistics given in this article is misleading. Randomisation is a major part of traditional sta
by conjectures 7y ago
I mostly agree with Pearl.
However, the account of traditional statistics given in this article is misleading. Randomisation is a major part of traditional stats, and it is inherently a causal hypothesis: breaking the links between unobserved covariates and treatment regimes.
An important alternate contemporary causal inference framework by Rubin has origins in a 1923 thesis...
...but the content of Pearl's approach seems superior; if you ignore the academic spats.
- fluentmundo 7y ago>> Randomisation is a major part of traditional stats, and it is inherently a causal hypothesis Yes, randomization is central to classical statistics, but no, it is not inherently causal. Drawing a random sample from a bivariate distribution (X,Y) is key to doing a lot (though not all) of classical statistical inference (think of estimating slopes in regression), but the randomization does not imply anything about the causal relationship between X and Y. When you speak of randomization in the context of "treatment regimes," you are thinking about randomized controlled trials, which the piece does analyze explicitly, in some detail. So in this sense the account given in the essay is not misleading.
- conjectures 7y agoI'm afraid you're mistaken. Randomisation allows one to make the strong causal assumption that the treatment regime allocation is unrelated to any of the other variables, observed or unobserved. Anyway, the section you're pointing to agrees with me. It just happens to be overlooked when they summarise...
- fluentmundo 7y agoNo, I am not mistaken; you are confused. And your confusion is very pervasive in the technical community. You're not talking about randomization in general when you speak of a "treatment regime." You're talking about randomization in a causal experiment such as a randomized controlled trial. But "randomization" is a broader thing than randomization in a controlled experiment. The assumption of randomization is made for almost all classical statistical inference, which has nothing at all to do with causation. Say you want to do basic linear regression: you want to estimate the slope for Y regressed on X. The most stringent form of inference works like this: you draw a random sample (X_i, Y_i), modeled as n independent and identically distributed realizations from the joint distribution (X,Y). Etc. This is certainly a stochastic model; we require randomization (or some approximation of it) to do inference. But it has nothing whatsoever to do with causality.