6 ms·
Seven basic rules for causal inference
- lordnacho 2y agoThis is brilliant. The whole causal inference thing is something I only came across after university, either I missed it or it is a hole in the curriculum, because it seems incredibly fundamental to our understanding of the world. The thing that made be read into it was a quite interesting sentence from lesswrong, saying that actually the common idea that correlation does not imply causation is wrong. Now it's not wrong in the face-value sense, it's wrong in the sense that actually you can use correlations to learn something about causation, and there turns out to be a whole field of study here.
- Vecr 2y agoWhen did you go to university? The terminology here came from Pearl 2000, and it probably took years and years after that to diffuse out.
- lordnacho 2y agoI thought Pearl was writing from 1984 onwards? I was at university around the millennium.
- janto 2y agoCausality (2000) made the topic accessible (to students and lecturers) as a single book.
- mister_jervis 2y ago[flagged]
- jerf 2y ago"correlation does not imply causation is wrong" That's a specific instance of a more general problem in the "logical fallacies", which is that most of them are written to be true in an absolutist, Aristotelian frame. It is true that if two things are correlated you can not therefore infer a rigidly 100% chance that there is a causative relationship there. And that's how Aristotelian logic works; everything is either True or False and if there is anything else it is as most "Indeterminate" and there is absolutely, positively, no in betweens or probabilities or anything else. However, consider the canonical "logical fallacy": 1. A -> B. 2. B 3. Therefore, A. It is absolutely a logical fallacy in the Aristotelian sense. Just because B is there does not mean A is. However, probabilistically, if you are uncertain about A, the presence of B can be used to update your expected probability of A. After all, this is exactly what Bayes' rule is for! Many of the "fallacies" can be rewritten to be useful probabilistically, and aren't quite as fallacious as their many internet devotees fancy. It is certainly reasonable to be "suspicious" about correlations. There often is a "there" there. Of course, whether you can ever figure out what the "there" is is quite a different question; https://gwern.net/everything https://gwern.net/everything really gets in your way. (I also recommend https://gwern.net/causality https://gwern.net/causality ). The upshot is basically 1. the glib dismissal that correlation != causation is, well, too glib and throws away too many things but 2. it is still true you still generally can't assume it either. The reality of the situation is exceedingly complicated.
- boxfire 2y agoI liked the way Pearl phrased it originally. A calculus of anti-correlations implies causation. That makes the nature of the analysis clear and doesn't set of the classic minds alarm bells.
- cubefox 2y agoUnfortunately this calculus is exceedingly complicated and I haven't even seen a definition of "a causes b" in terms of this calculus. One problem is that Pearl and others make use of the notion of "d-separation". This allows for elegant proofs but is hard to understand. I once found a paper which replaced d-separation with equivalent but more intuitive assumptions about common causes, but I since forgot the source. By the way, there is also an alternative to causal graphs, namely "finite factored sets" by Scott Garrabrant. Probably more alternatives exist. Though I don't know more about (dis)advantages.
- geye1234 2y agoI don't disagree with the substance of your comment, but want to clarify something. Lesswrong promulgated a seriously misleading view of Aristole as some fussy logician who never observed reality and was unaware of probability, chance, the unknown, and so on. It is entirely false. Aristotle repeats, again and again and again, that we can only seek the degree of certainty that is appropriate for a given subject matter. In the Ethics, perhaps his most-read work, he says this, or something like it, at least five times. I mention this because your association of the words "absolutist" and "Aristotelian" suggests your comment may have been influenced by this. ISTM that there are two entirely different discussions taking place here, not opposed to each other. "Aristotelian" logic tends to be more concerned with ontology -- measles causes spots, therefore if he has measles, then he will have spots. Whereas the question of probability is entirely epistemological -- we know he has spots, which may indicate he has measles, but given everything else we know about his history and situation this seems unlikely; let's investigate further. Both describe reality, and both are useful. So the fallacies are entirely fallacious: I don't think your point gainsays this. But I agree that, to us, B may suggest A, and it is then that the question of probability comes into play. Aquinas, who was obviously greatly influenced by Aristotle, makes a similar point somewhere IIRC (I think in SCG when he's explaining why the ontological argument for God's existence fails), so it's not as if this is a new discovery.
- currymj 2y agoRigorous causal inference methods are just now starting to diffuse into the undergraduate curriculum, after gradually becoming part of the mainstream in a lot of social science fields. But this is just happening. Judea Pearl is in some respects a little grandiose, but I think he is right to be express shock that it took almost a century to develop to this point, given how long the basic tools of probability and statistics have been fairly mature.
- currymj 2y agoRule 2 (“causation creates correlation”) would be strongly disputed by a lot of people. It relies on the assumption of “faithfulness” which is not discussed until the bottom of the article. This is a very innocent sounding assumption but it’s actually quite strong. In particular it may be violated when there are control systems or strategic agents as part of the system you want to study — which is often the case for causal inference. In such scenarios (eg the famous thermostat example) you could have strong causal links which are invisible in the data.
- apwheele 2y agoThis was my thought as well. I don't like showing the scatterplots in these examples, as "correlation" I think is more associated with the correlation coefficient than the more generic independence that the author means in this scenario. E.g. a U shape in the scatterplot may have a zero correlation coefficient but is not conditionally independent.
- bdjsiqoocwk 2y ago> E.g. a U shape in the scatterplot may have a zero correlation coefficient but is not conditionally independent. Ok this is correct, but has nothing to do with causality. Whether or not two variables are correlated and whether or not they are independent, and when one does or doesn't imply the other, is a conversation that can be had without resorting to the concept of causality at all. And in fact that's how the subject is taught at an introductory level basically 100% of the times.
- cubefox 2y ago> Ok this is correct, but has nothing to do with causality. It does. Dependence and independence have a lot to do with causation, as the article explains. > Whether or not two variables are correlated and whether or not they are independent, and when one does or doesn't imply the other, is a conversation that can be had without resorting to the concept of causality at all. Yes, but this is irrelevant. It's like saying "whether or not someone is married is a conversation that can be had without resorting to the concept of a bachelor at all". You can talk about (in)dependence without talking about causation, but you can't talk in detail about causation without talking about (in)dependence.
- Vecr 2y agoAre the assumptions "No spurious correlation", "Consistency", and "Exchangeability" ever actually true? If a dataset's big enough you should generally be able to find at least one weird correlation, and the others are limits of doing statistics in the real world.
- levocardia 2y agoSome situations guarantee certain assumptions: Randomization, for example, guarantees exchangeability.
- shiandow 2y agoThis is missing my favourite rule. 0. The directions of all arrows not part of a collider are statistically meaningless.
- Vecr 2y agoWhat's not part of a collider? Good luck with your memory in that case.
- 082349872349872 2y agoI'm guessing they mean that given a bunch of correlated nodes but no collider (in which case the casual graph must be a tree of some sort) you not only don't know if the tree be bushy or linear, you don't even know which node may be the root. (bushy trees, of which there are very many compared with linear ones, would be an instance of Gwern's model* of confounds being [much] more common than causality?) * https://news.ycombinator.com/item?id=41291636 https://news.ycombinator.com/item?id=41291636
- Vecr 2y agoRight, but your memory functions as a collider, if there are literally no colliders anywhere you by definition won't be able to remember anything.
- dkga 2y agoI highly suggest this paper here for a more complete view of causality that nests do-calculus (at least in economics): Heckman, JJ and Pinto, R. (2024): “Econometric causality: The central role of thought experiments”, Journal of Econometrics, v.243, n.1-2.
- fn-mote 2y agoWhy should you look this paper up? It argues that certain approaches from statistics and computer science are limited, and (essentially) that economists have a better approach. YMMV, but the criticisms are specific (whether or not you buy the "fix"). From the paper: > Each of the recent approaches holds value for limited classes of problems. [...] The danger lies in the sole reliance on these tools, which eliminates serious consideration of important policy and interpretation questions. We highlight the flexibility and adaptability of the econometric approach to causality, contrasting it with the limitations of other causal frameworks.
- deleted 2y ago[deleted]
- Rhapso 2y agoI'm keeping this link, taking a backup and handing it out whenever i can. It is succinct and effective. These are concepts i find myself constantly having to explain and teach and they are critical to problem solving.
- 082349872349872 2y agoCan these seven be reduced to three basic rules? - controlling for a node increases correlation among pairs where both are ancestors - controlling for a node does not affect (the lack of) correlation among pairs where at least one is categorically unrelated (shares no ancestry with that node) - controlling for a node decreases correlation among pairs where both are related but at least one is not an ancestor
- raymondh 2y agoIs there a simple R example for Rule 4?
- elsherbini 2y agoIt is sort of tautological: # variable A has three causes: C1,C2,C3 C1 <- rnorm(100) C2 <- rnorm(100) C3 <- rnorm(100) A <- ifelse(C1 + C2 + C3 > 1, 1, 0) cor(A, C1) cor(A, C2) cor(A, C3) # If we set the values of A ourselves... A <- sample(c(1,0), 100, replace=TRUE) # then A no longer has correlation with its natural causes cor(A, C1) cor(A, C2) cor(A, C3)
- abeppu 2y agoAt the bottom, the author mentions that by "correlation" they don't mean "linear correlation", but all their diagrams show the presence or absence of a clear linear correlation, and code examples use linear functions of random variables. They offhandedly say that "correlation" means "association" or "mutual information", so why not just do the whole post in terms of mutual information? I think the main issue with that is just that some of these points become tautologies -- e.g. the first point, "independent variables have zero mutual information" ends up being just one implication of the definition of mutual information.
- jdhwosnhw 2y agoThis isnt a correction to your post, but a clarification for other readers: correlation implies dependence, but dependence does not imply correlation. Conversely, two variables share non-zero mutual information if and only if they are dependent.
- islewis 2y agoCould you give some examples of dependence without correlation?
- xtacy 2y agoYou can check the example described here: https://stats.stackexchange.com/questions/644280/stable-violation-of-faithfulness https://stats.stackexchange.com/questions/644280/stable-viol... Judea Pearl’s book also goes into the above in some detail, as to why faithfulness might be a reasonable assumption.
- abeppu 2y agoA clear graphical set of illustrations is the bottom row in this famous set: https://en.wikipedia.org/wiki/Correlation#/media/File:Correlation_examples2.svg https://en.wikipedia.org/wiki/Correlation#/media/File:Correl... They have clear dependence; if you imagine fixing ("conditioning") x at a particular value and looking at the distribution of y at that value, it's different from the overall distribution of y (and vice versa). But the familiar linear correlation coefficient wouldn't indicate anything about this relationship.
- levocardia 2y ago>Controlling for a collider leads to correlation This is a big one that most people are not aware of. Quite often, in economics, medicine, and epidemiology, you'll see researchers adjust for everything in their regression model: income, physical activity, education, alcohol consumption, BMI, ... without realizing that they could easily be inducing collider bias. A much better, but rare, approach is to sit down with some subject matter experts and draft up a DAG - directed acyclic graph - that makes your assumptions about the causal structure of the problem explicit. Then determine what needs to be adjusted for in order to get a causal estimate of the effect. When you're explicit about your causal assumptions, it makes it easier for other researchers to propose different causal structures, and see if your results still hold up under alternative causal structures. The DAGitty tool [1] has some cool examples. [1] https://www.dagitty.net/dags.html https://www.dagitty.net/dags.html
- kyllo 2y agoCollider bias or "Berkson's Paradox" is a fun one, there lots of examples of it in everyday life: https://en.wikipedia.org/wiki/Berkson%27s_paradox https://en.wikipedia.org/wiki/Berkson%27s_paradox
- chrsig 2y ago> Rule 8: Controlling for a causal descendant (partially) controls for the ancestor perhaps this is a quaint or wildly off base question, but an honest one, please forgive any ignorance: Isn't this essentiallydefining the partial derivative? Should one arrive at the calculus definition of a partial derivative by following this?
- bubblyworld 2y agoYou probably could if you interpret that sentence very creatively. But I think it's useful to remember that this is mathematics, and words like "control", "descendant" and "ancestor" have specific technical meanings (all defined in the article, I believe). The technical meaning of that sentence has to do with probability theory (probability distributions, correlation, conditionals), and not so much calculus (differentiable functions, limits, continuity).
- nomilk 2y agoHumble reminder of how easy R is to use. Download and install R for your operating system: https://cran.r-project.org/bin/ https://cran.r-project.org/bin/ Start it in the terminal by typing: R Copy/paste the code from the article to see it run!
- curiousgal 2y agoCan't use R without RStudio. It so much better than the terminal.
- nomilk 2y agoAgree RStudio makes R a dream, but isn't necessary for someone to run the code in the article =)
- throwway_278314 2y agoreally??? I've developed in R for over a decade using two terminal windows. One runs vim, the other runs R. Keyboard shortcuts to send R code from vim to R. first google hit if you want to try this yourself: https://www.freecodecamp.org/news/turning-vim-into-an-r-ide-cd9602e8c217/ https://www.freecodecamp.org/news/turning-vim-into-an-r-ide-... Sooooooo much better than "notebooks". Hating on "notebooks" today.
- carlmr 2y ago>Humble reminder of how easy R is to use. I had to learn R for a statistics course. This was a long time ago. But coming from a programming background I never found any other mainstream language as hard to grok as R. Has this become better? Is it just me that doesn't get it?
- incognito124 2y agoR is my least favorite language to use, thanks to the uni courses that force it https://github.com/ReeceGoding/Frustration-One-Year-With-R https://github.com/ReeceGoding/Frustration-One-Year-With-R
- crystal_revenge 2y ago> Independent variables are not correlated But it's important to remember that dependent variables can also be not correlated. That is no correlation does not imply independence. Consider this trivial case: X ~ Uniform(-1,1) Y = X^2 Cor(X,Y) = 0 Despite the fact that Y's value is absolutely determined by the value of X.
- TheRealPomax 2y agoThis is also why it's important to look at your plots. Because simply looking at your scatter plot makes it really obvious what methods you can't use, even if it doesn't really tell you anything about what you should use.
- antognini 2y agoThe author is using "correlation" in a somewhat non-standard way. He isn't referring to linear correlation as you are, but any sort of nonzero mutual information between the two variables. So in his usage those two variables are "correlated" in your example.
- arunsupe 2y agoGreat post. It's nice that these rules can be trivially demonstrated by simulation. The simulation (and visuals) helps validate the concepts.