13 ms·
Oh what fun to discover the horror of causality! For some areas of research, truly understanding causality is essentially impossible - if well-controlled exper
by uniqueuid 2y ago
Oh what fun to discover the horror of causality!
For some areas of research, truly understanding causality is essentially impossible - if well-controlled experiments are impossible and the list of possible colliders and confounders is unknowable.
The key problem is that any causal relation can be an illusion caused by some other, unobserved relation!
This means that in order to show fully valid causal effect estimates, we need to
- measure precisely
- measure all relevant variables
- actively NOT measure all harmful (i.e. falsely correlated) variables
I heartily recommend the book of why [1] by Pearl and Mackenzie for a deeper reading and the "haunted DAG" in McElreath's wonderful Statistical Rethinking.
[1] https://en.wikipedia.org/wiki/The_Book_of_Why https://en.wikipedia.org/wiki/The_Book_of_Why
- kqr 2y agoPearl's Causality is very high on my "re-read while making flashcards" list. It is depressing how hard it is to establish causality, but also inspiring how causality can be teased out of observational statistics provided one dares assume a model on which variables and correlations are meaningful.
- uniqueuid 2y ago"provided one dares assume ..." - that's a great quote which I'll steal in the future if you allow! Most things we learn about DAGs and causality are frustrating, but simulating a DAG (e.g. with lavaan in R) is a technique that actually helps in understanding when and how those assumptions make sense. That's (to me) a key part of making causality productive.
- KempyKolibri 2y agoI’ve heard Miguel Hernán’s “What If” is also excellent, but not got round to reading it.
- uniqueuid 2y agoYes it's great! There is also this great book on causality in ML, but it's a much heavier read: Chernozhukov, V., Hansen, C., Kallus, N., Spindler, M., & Syrgkanis, V. (2025). Causal Inference with ML and AI.
- levocardia 2y agoFor a lighter introduction to Hernán’s ideas check out: "The C-Word: Scientific Euphemisms Do Not Improve Causal Inference From Observational Data" (https://pmc.ncbi.nlm.nih.gov/articles/PMC5888052/ https://pmc.ncbi.nlm.nih.gov/articles/PMC5888052/) "Does water kill? A call for less casual causal inferences" (https://pmc.ncbi.nlm.nih.gov/articles/PMC5207342/ https://pmc.ncbi.nlm.nih.gov/articles/PMC5207342/)
- alexpetralia 2y agoI have reflected on a good definition of causality and would be curious if anyone has thoughts or critiques of it. I am repasting part of my essay below. (https://alexpetralia.com/2023/02/25/statistics-only-gives-correlations-so-what-about-causation-part-11/ https://alexpetralia.com/2023/02/25/statistics-only-gives-co...) -- Can we nevertheless extract causality from correlation? I would argue that, theoretically, we cannot. Practically speaking, however, we frequently settle for “very, very convincing correlations” as indicative of causation. A correlation may be persuasively described as causation if three conditions are met: Completeness: The association itself (R²) is 100%. When we observe X, we always observe Y. No bias: The association between X and Y is not affected by a third, omitted variable, Z. Temporality: X temporally precedes Y.
- kqr 2y agoI feel like you have this backwards. In the assignment Y:=2X, each unit of Y is caused by half a unit of X. In the game where we flip a coin at fair odds, if you have increased your wealth by 8× in 3 tosses, that was caused by you getting heads every toss. Theoretically establishing causality is trivial. The problem comes when we try to do so practically, because reality is full of surprising detail. > No bias: The association between X and Y is not affected by a third, omitted variable, Z. This is, practically speaking, the difficult condition. I'm not so convinced the others are necessary (practically speaking, anyway) but you should read Pearl if you're into this!
- uniqueuid 2y agoYou are missing one crucial additional condition: - No colliders have been included in the analysis, which would introduce appearance of causality that does not exist
- HPsquared 2y agoRuling out all Z is the almost-impossible part. It's hard to prove a negative, especially with incomplete information.
- stonemetal12 2y agoWhat of the double slit experiment, where observation changes the outcome? Do we call observation the cause of the outcome?
- currymj 2y agoeven if you hit all the assumptions you need to make Pearl/Rubin causality work, and there is no unobserved factor to cause problems, there is still a philosophical problem. it all assumes you can divide the world cleanly into variables that can be the nodes of your DAG. The philosopher Nancy Cartwright talks about this a lot, but it’s also a practical problem.
- deleted 2y ago[deleted]
- shadowgovt 2y agoAnd this is even before we get into the philosophical / epistemological questions about "cause." You can make the argument, from correlative data, that bridges and train tracks cause truck accidents. And more importantly, if you act like they do when designing roadways, you actually will decrease truck accidents. But it's a common-sense-odd meaning of causality to claim a stationary object is acting upon a mobile object...
- QuantumGood 2y agoThat colliders and confounders have technical definitions is not known by some: ------------------ Confounders ------------------ A variable that affects both the exposure and the outcome. It is a common cause of both variables. Role: Confounders can create a spurious association between the exposure and outcome if not properly controlled for. They are typically addressed by controlling for them in statistical models, such as regression analysis, to reduce bias and estimate the true causal effect. Example: Age is a common confounder in many studies because it can affect both the exposure (e.g., smoking) and the outcome (e.g., lung cancer). ------------------ Colliders ------------------ A variable that is causally influenced by two or more other variables. In graphical models, it is represented as a node where the arrowheads from these variables "collide." Role: Colliders do not inherently create an association between the variables that influence them. However, conditioning on a collider (e.g., through stratification or regression) can introduce a non-causal association between these variables, leading to collider bias. Example: If both smoking and lung cancer affect quality of life, quality of life is a collider. Conditioning on quality of life could create a biased association between smoking and lung cancer. ------------------ Differences ------------------ Direction of Causality: Confounders cause both the exposure and the outcome, while colliders are caused by both the exposure and the outcome. Statistical Handling: Confounders should be controlled for to reduce bias, whereas controlling for colliders can introduce bias. Graphical Representation: In Directed Acyclic Graphs (DAGs), confounders have arrows pointing away from them to both the exposure and outcome, while colliders have arrows pointing towards them from both the exposure and outcome. ------------------ Managing ------------------ Directed Acyclic Graphs (DAGs): These are useful tools for identifying and distinguishing between confounders and colliders. They help in understanding the causal structure of the variables involved. Statistical Methods: For confounders, methods like regression analysis are effective for controlling their effects. For colliders, avoiding conditioning on them is crucial to prevent collider bias.
- lenzm 2y agoIf you have to start with apologies then you know, just stop and don't post.
- 2y ago
- dan_mctree 2y agoAnd even if you do know there's causality (eg: the input variable X is part of software that provides some output Y), the exact nature of the causality can be too complex to analyze due to emergent and chaotic effects. It's seldom as simple as: an increase in X will result in an increase in Y