16 ms·
Correlation is usually not causation. But why not?
- zdw 12y agoBecause you can find things that are correlated, but have no practical relationship - many (frequently humorous) examples here: http://tylervigen.com http://tylervigen.com
- ekianjo 12y agoExcellent link, thanks for posting :)
- gwern 12y agoDatamining correlations like that confuses two issues: sampling error and 'everything is correlated'. When you find a correlation like that, it's unclear whether it's a spurious correlation which will disappear if you just collect more data (and caused by slicing the data in a thousand ways without any compensation by this) or whether there genuinely is some sort of opaque correlation which will persist no matter how much data you collect (maybe 'US spending on science, space, and technology' really does correlate with 'Suicides by hanging, strangulation and suffocation' through something like government stimuluses and economic growth).
- api 12y agoPot smoking kills bees is my favorite... ought to make it into a t-shirt.
- fizixer 12y agoDisagree with the title. Correlation does imply casuation a lot of the times (especially for simple systems). But not always. Therefore, the caution is not to assume it apriori, but pursue further investigation to confirm or reject it. Even when it is rejected, a lot of those cases result in a third variable being the cause behind the correlated "effect" variables.
- jerf 12y ago"Disagree with the title." Perhaps you should read the article before posting your disagreement. (As well as the several other people who appear to have paragraph-sized responses to 8 words, rather than the actual article.)
- lotsofpulp 12y agoThe problem with using "correlation does not imply causation" outside of a mathematical context is that in math, saying A implies B means that if condition A is true, then condition B is true. In common vernacular, implies means to suggest, hence the confusion.
- fizixer 12y ago12 days late but still: I just realized "Correlation does imply casuation a lot of the times" is a completely meaningless statement that I made (if correlation is not because of causation even some of the times then that's simply 'correlation doesn't imply causation'). What I really had in mind (which I think is correct) is that 'correlation ends up being due to causation a lot of the times'.
- jostmey 12y agoStatistics serves as a tool to overcome our cognitive biases. But what if these biases are at the center of learning? Take for example the Gambler's Fallacy where a player believes she can predict the outcome of a coin toss with greater certainty than is possible. Obviously she cannot. But if I had to design a Machine Learning algorithm, I would certainly want it to always assume that a pattern existed. That way, if the data were predictable, the algorithm would be able to take advantage of it.
- eternauta3k 12y agoA smarter algorithm would refrain from betting on the coin.
- jostmey 12y agoThat's great if you know a priori that you cannot predict the outcome of what you are betting with.
- nitrogen 12y agoI don't think it would be necessary to know a priori that something is unpredictable; failing to find a pattern after some number of observations should allow an algorithm to determine that the event is not predictable by that algorithm.
- baddox 12y agoThe trouble is that you can't reliably distinguish random data (or data determined by factors you're not considering) from patterned data.
- thaumasiotes 12y agoHow do you think random number generators are tested?
- D_Alex 12y ago
- QuantumChaos 12y agoCorrelation vs causation is not actually a complex mathematical problem. The issue is more philosophical. Bayesian networks are a highly sophisticated and flexible framework for thinking about causation. And yet the essence of Bayesian networks can be captured in much simpler methods like Instrumental Variables or simply regressions with controls. In all cases, the true distribution of observables (which we can estimate from samples) combined with some assumptions about the possible nature of causality (either very sophisticated in the case of Bayesian networks, or very simple in the case of instrumental variables) lead to the actual causal relationships. Where do these assumptions come from? Consider a randomized trial. Even though physically, it is possible that some hidden cause influenced both the random number generator which selected patients into a trial, and also whether patients will get better. And yet people universally believe that there is no such mechanism, and so they believe that randomized trials prove the causal relationship between taking a drug and getting better. No physicist ever wrote down an equation proving this. It is simply something we deduced from our human understanding of how nature works. Put simply, causality is not a physical notion, neither is it a statistical notion. It is a part of our intuitive understanding of the physical world, in which a higher level notion of causality exists, beyond that described by special relativity.
- pasharayan 12y agoTo add to your point, we cannot prove or disprove the existence of causation (Try to conceive a falsifiable experiment about causation and I would argue you would end up with a metaphysical crisis). The 'is-ought' issue David Hume showed, where just because something 'is' a way does not indicate how it 'ought' to be or will continue to be, has highlighted to us how difficult it is to think this problem for hundreds of years. Another way to look at it is to ask "Who or what is to guarantee the law of physics will remain the same tomorrow? What is to stop the speed of light changing to 1 mile per hour?" This isn't to say the article posted is of no use. Having a 'graph-like' mental model of how things work is incredibly useful, as most education simplifies real-world problems into a few key issues. Although most non-computer science issues can be reduced successfully using the 80/20 rule in real life, sometimes some problems require us to look at the 100s of contributing factors to allow us to solve the problem we're facing properly. The more 'graph-like' problem solving becomes acceptable as way to solve issues the better of everyone will be.
- jrapdx3 12y agoThis question is bound to be a can of worms. There has been a great deal written about the matter, particularly in regard to observational studies. Drawing causal inference is overwhelmingly likely to be wrong when there is a good chance that unknown variables are influencing the correlates observed. In health-related sciences that more often than not is the case. A few years ago there was a study correlating hours of TV watched and ADHD symptoms in children. The news media picked up on these "findings" and of course the causal influence of TV on ADHD was reported. It was obvious that saying watching TV caused ADHD was absurd, that other variables weren't taken into account, e.g., some other characteristic of ADHD kids prompted watching TV more than other kids. There was a great article published in PLOS several years ago (ATM I don't have the link) showing mathematically that the odds were about 1 in a million that an observational study like the above would turn out to be a "true" causal relationship, and the author concluded most published studies were junk. In experimental studies, variables are limited and controlled as well as possible. With fewer and known variables, correlations would have a greater chance of revealing a reproducible causal relationship among events. The discussion gets tripped up when it comes to defining "cause" or "causal relationship". The theory is controlled as an experiment may be, there's a possibility that unknown variables were present and affected the phenomena occurring in the experiment. Conclusions can't be absolute, but only true to some probability. I think the history of science over the last 100 years or so has something to say about the nature of "causal relationships".
- nmjohn 12y ago> There was a great article published in PLOS several years ago (ATM I don't have the link) showing mathematically that the odds were about 1 in a million that an observational study like the above would turn out to be a "true" causal relationship, and the author concluded most published studies were junk. I very much doubt the second part of that, that most published studies are junk. The idea that not showing a causal link and only a correlational one makes a study junk is not held by anyone in the field who garners a whole lot of respect or notoriety. I'm not saying the quality is equal whatsoever, simply that correlation studies are not inherently junk - some are fantastic and some are quite the opposite. There is no question that journals in general are pumping out a significant amount of junk (in studies of all types), but I speculate the root cause of that has more to do with the rise of "publish or perish" and significant increases in grad school enrollment. And even worse the notion that for a grad student, the number of publications they have their name attributed to is more important than the content they publish in terms of employment after they graduate. So there is a situation where people have more pressure to publish than ever before [0], there is more competition for scarce funding so studies are vastly underfunded and studies are rushed so another can begin to add to the resume. That doesn't mean there is less good science being done either! There probably is more good science being published now than ever before, the problem is the signal:noise ratio has gone down making it harder for good studies to get media attention, and easier for the media to latch on to whatever story they think will get viewers. [0]: Sidenote, this also makes it increasingly challenging not only to have high quality research, but research that only is making a correlation as opposed to going through and providing evidence to claim you might have a causal link.
- deleted 12y ago[deleted]
- tel 12y agoThe author more than understands this. In the first two paragraphs he address the much more interesting question about whether, in realistic causal DAGs, potential correlations grow at a faster rate than actual causal links—the idea being that if they did it would harm the notion that seeing a correlation improves the hope of causation due to pigeonholing. The author doesn't assert that this is true, but it definitely means he's thinking about this quite hard.
- deleted 12y ago[deleted]
- dj-wonk 12y agoThis article starts off with a common mistake. It is ok to create the 3 categories: > If I suspect that A→B, and I collect data and establish beyond doubt that A&B correlates r=0.7, how much evidence do I have that A→B? > you can divvy up the possibilities as: 1. A causes B 2. B causes A 3. both A and B are caused by a C So far so good, but here is the problem: > Even if we were guessing at random, you’d expect us to be right (at at least 33% of the time... No. It is not valid to assume each possibility is equally likely. If you do so, you are bringing your own assumptions to the problem. If you ever find yourself assuming a distribution, pause and consider testing your assumption.
- Jemaclus 12y agoIf you read the whole article, he addresses this about halfway down.
- dj-wonk 12y ago@Jemaclus I'm not sure why you assumed that I did not read the article. Maybe you disagreed with my conclusion and assumed I had not read the article? To your question, I would ask: Where, exactly, does the author "address this"? (Maybe we are talking about different meaning of "this"? I explain the fallacy I'm talking about in a longer comment on a sister thread (by "sister" I mean up one level, over one, down one). Did you mean this part? > It turns out, we weren’t supposed to be reasoning ‘there are 3 categories of possible relationships, so we start with 33%’, but rather: ‘there is only one explanation “A causes B”, only one explanation “B causes A”, but there are many explanations of the form “C1 causes A and B”, “C2 causes A and B”, “C3 causes A and B”…’, and the more nodes in a field’s true causal networks (psychology or biology vs physics, say), the bigger this last category will be. This is also fallacious, in my opinion. See my longer comment. In short, if you "count up" categories in this way, you are assuming information that you don't have; doing so is a matter of belief, not analysis. Or did you mean this part? > they might not be reasoning in a causal-net framework at all, but starting from the naive 33% base-rate you get when you treat all 3 kinds of causal relationships equally. > This could be shown by eliciting estimates and seeing whether the estimates tend to look like base rates of 33% and modifications thereof. This chunk of text does not debunk the claim that a 33% base rate is reasonable as a starting point. My point is that the 33% starting point was not reasonable in the first place.
- graycat 12y agoBad winter weather can cause auto accidents, and we expect a positive correlation between bad winter storms and winter auto accidents. Okay, but in the northern hemisphere, living in more northern latitudes also correlates with winter auto accidents but does not cause them. For heart disease, we know that the main causes have to do with aging. Well, then, since now the audience for TV news is comparatively old, we can expect that watching TV news has positive correlation with heart disease. Still watching TV news does not cause heart disease. In the US NE, hurricanes are positively correlated with pretty leaves on the trees, but the leaves do not cause the hurricanes, and the hurricanes do not cause the colors in the leaves. Instead, hurricanes are caused by the surface waters of the Atlantic Ocean hot from summer sun with cooler air on top, and that situation is caused by the fall weather with also causes the colored leaves. So, colored leaves have a spurious correlation with hurricanes but do not cause them. Correlation is much more common and much easier to establish than causality. Usually the convincing evidence of causality if some basic physical connection.
- coldtea 12y ago>Well, then, since now the audience for TV news is comparatively old, we can expect that watching TV news has positive correlation with heart disease. Still watching TV news does not cause heart disease. Well, I'm not so sure. All this sitting to watch TV news, plus all the stress from bad news and fear-mongering...
- TheSpiceIsLife 12y agoWatching (bad) fear monger news doesn't cause stress. Rather, the story you tell yourself about what the news means causes the stress. One person (say, me for example) can watch the (bad) news and have a good laugh (call that a , while another (one of my housemates, for example) will watch the same thing and have a negative response to it.
- coldtea 12y agoWell, that's like "gun's don't kill people, bullets (or holes in vital organs) kill people". Technically correct, but doesn't really take away from the first thing.
- fiatmoney 12y agoIn reality, causal mechanisms for the phenomena we most care about (economic and social phenomena particularly) are so fantastically complicated that you never really "figure it out". Nevertheless, if you can exploit the hypothesized causation to optimize some function, you've won, regardless of the metaphysics. Even in the context of physical processes, like the biological processes Gwern mentions, the notion of "establishing causality" is much more of a regulatory artifact than anything else. You might call this the machine learning approach (optimize some objective function, regardless of mechanism) as opposed to the statistics approach (generate some measured claim about a process itself).
- ghshephard 12y agoThe article is a little dense for me, and presumes knowledge about Probabilistic Graphical Models (PGMs) and directed acyclic graphs / causal Bayesian networks (DAGs) without introducing the background knowledge - so, I expect my reading of it missed a lot of the details. But the one thing I came out wondering, is if you do a proper randomized double-blinded study with a large population sample, say, taking 1000 people with a particular illness, completely randomizing them, and then giving 500 of the sample a particular treatment, and 500 of the sample a placebo that is indistinguishable from the treatment. If an unbiased third-party, then observes the groups (still without knowledge of which group had which treatment), and if one group shows markedly different results (say, 490/500 of the treatment group recover, and only 10/500 placebo group recover), is it fair to say that in this particular scenario, that correlation of recovery with the treatment implied that the treatment caused the recovery?
- mjmahone17 12y agoYes. In this case, you have two outcomes: Recovering and Not Recovering. You also have two starting states: Drug and Placebo (which both came off of the original starting state, Participant in Study, with a 50% probability). So your Drug -> Recover has a 98% probability, vs. Drug -> Not Recover of 2%. Likewise, Placebo -> Recover has only a 2% probability, vs. Placebo -> Not Recover of 98%. In this case, you know with very high probability that the fork you stuck in the road caused two very clear, disparate outcomes. Especially because you actually began the "give drug" phase. However, usually medical trials have results more along the lines of: Placebo group of 35 participants: 10 recovered in < 5 days, 12 recovered in < 10 days, and 13 failed to recover. Drug group of 35 participants: 17 recovered in < 5 days, 13 recovered in < 10 days, and 5 failed to recover. With data like that, clearly SOMETHING causes people to recover, and clearly your drug is not the underlying cause of recovery. It may be supporting whatever process is causing recovery, but it's very unclear whether that's even the case.
- aaron695 12y agoWell it's fair to say that experiment shows causation but no I don't think you can say the correlation shows it works. I think it goes against the definition of the word..... I found the article too dense to get as well, but I do think 'perhaps' we could use correlation more to assume causation. I think we are too cautious and certainly the haters always bring in correlation to stop science articles they think don't follow their beliefs. Not sure if that's compatible with the OP or not, it hurt my head.
- dusklight 12y agoAll I know about causation and correlation I learnt hunting bugs in large legacy software systems. In that environment I got the impression that correlation almost never equalled causation, but that's only because the hardest bugs, the ones I remembered, were hard because the obvious correlations did not help identify the root cause. A similar argument might be made for scientific studies: most of the easy causes that can be identified from correlation have already been found, leaving the majority of new studies with correlations that don't easily establish causation.
- api 12y agoI think this low-hanging-fruit idea is generally true for all of science: after centuries of science as a profession, pretty much all the easy stuff has already been done. It's actually an argument for more cojones in science-- being willing to do bold stuff and explore "crazy" hypotheses. If all the easy stuff is done already, then picking methodically and timidly among the dregs is unlikely to ever yield anything.
- mattfenwick 12y agoAn alternative viewpoint: everything's easy after it's been done, but it's hard up until then. > after centuries of science as a profession, pretty much all the easy stuff has already been done. But with the benefit of all that has already been done, shouldn't we be able to do things now that were previously impossible? In other words, "easy in 2014" != "easy in 1200". > It's actually an argument for more cojones in science-- being willing to do bold stuff and explore "crazy" hypotheses. Maybe doing whatever's easy is the most efficient (by time, money) way to make progress?
- api 12y agoI'm not sure... seems like what you say may hold for a while, but eventually you start hitting a more objective sort of hard: things that are hard for human beings to comprehend due to the limitations of our intelligence itself. Beyond that there's probably an even harder hard-- when you start actually running out of new things to discover. Can there really be an infinite number of physical laws, principles, and useful relations? Or at some point have you actually found most of basic physics? Once you start hitting those, you've either entered a permanent era of diminishing returns or one where you can only really make progress by radically redefining problems, making leaps, or trying wild and crazy ideas in the hopes of unlocking some isolated seam of high-value research that isn't connected to the others in the fitness/value state space graph.
- sqrt17 12y agoNormally, causation has a couple confounders for correlation, such as - inverse causation - common cause - random correlations The first two should be relatively stable and reproducible, and we could then proceed to find out about causation with an intervention study (e.g. if "good education" directly causes "good job prospects", will it help if we give people good education that wouldn't normally get one? Or does it improve people's educational achievement when we give them better access to jobs?) The third isn't reproducible, but should normally be relatively rare. Why we're seeing more and more non-reproducible results is usually - People fishing around in data sets for correlations, or tweaking experiments until they find a correlation Because of this, we end up with a great deal of correlations that have nothing to do with causation. For a related but different perspective, see this article ("Language is never, ever random", Adam Kilgariff) http://www.kilgarriff.co.uk/Publications/2005-K-lineer.pdf http://www.kilgarriff.co.uk/Publications/2005-K-lineer.pdf
- JonnieCache 12y agoAnother cracking article from gwern. For those wondering, that snazzy looking `choose` combinatorial function is from clojure, and I assume other lisps: http://clojuredocs.org/incanter/incanter.core/choose http://clojuredocs.org/incanter/incanter.core/choose
- TeMPOraL 12y agoNo, it's just binomial coefficient. It's common to say "10 choose 2", it's the standard way of reading binomial coefficients out loud. See http://www.wolframalpha.com/input/?i=10+choose+2 http://www.wolframalpha.com/input/?i=10+choose+2.
- mgregory22 12y agoThe difference between correlation and causation is merely conventional. Does the striking of a match cause it to ignite? There is no way to prove that it does, only the correlation between the striking and the ignition makes us say they're causal. Correlation is the only way to determine what's causal and what's not. On the other hand, if you want to look at it philosophically, then the only sensible definition of "causation" is "anything necessary for a thing to exist." So what's necessary for a thing to exist? Every other thing in the universe that is not that thing! Pluto causes us to exist right now because it hasn't turned into a giant space goat and swallowed the Earth. Yes, the fact that something DIDN'T prevent the existence of a thing is also ultimately a cause of that thing's existence. Causation in its purest form can help us understand the nature of reality, but it can't help us predict anything. It's only when we draw a line between causes that we can control and causes that we can't that it becomes a practical tool (not that understanding reality isn't practical). So you can see, this whole correlation != causation discussion is pure nonsense. Get some philosophical skill before you try to discuss philosophical issues and stop embarrassing yourselves. Sheesh, it's ridiculous.
- return0 12y agoApart from semantic acrobatics, there is a practical need for defining causation in science and technology. The cause has to have a temporal precedence and be necessary. > Causation in its purest form can help us understand the nature of reality, but it can't help us predict anything huh?
- tlarkworthy 12y agoThe arrow of time. Randomised experiments.
- anon4 12y agoWhat? Your comment is pure nonsense. Of course you can prove that striking a match causes it to ignite. There's a long chain of events, each of which trivially physically verifiable, which starts with you moving the match head while applying a given amount of pressure against the material lining the matchbox, which then, due to friction, flakes off (both head and lining), the two react together to produce a spark and a flame starts. You can cut the chain down to micro-events and you'll find each one to be consistently verifiable. Or do an experiment - grab a mug and lift it. I claim that your arm motion while your hand had grabbed the mug caused it to ascend. We can again examine the basic physical events, cutting them as fine as you'd like, and we'll see that it really was you with your arm, who caused the mug to ascend. This is a cause-and-effect link -- when you can show how event A lead to event B via a chain of events, each of which directly causes the next. At the least you need to show that there is a sensible progression from the cause to the effect, in order to claim a causation. Correlation on the other hand is a description of past observations. You can observe that each time it snows people increase their energy expenditure in order to heat their homes. Does the increase in heating cause the snow? That doesn't seem likely, knowing how heating works. It could be that increased energy demands leads to more coal being burned, leading to larger clouds being formed and lowering the outside temperature. However, even though increased coal burning does lead to such effects, they're pretty small from just increased heating, so we can discount this hypothesis. On the other hand, you can show that cold weather causes snow. Water that falls in cold air crystallises and becomes snow. So there is a definite causative relationship between cold weather and snow. Could there also be a causative relationship between cold weather and heating? Well, assuming that people wish to live at a constant surrounding air temperature of about 20-30 degrees, that seems very likely. There are of course individuals who don't fit that profile, but they are very few (I can't actually provide a study which supports that claim, but just assume it for the purposes of this demonstration). In the end, we can say that while snow and heating are correlated, heating does not cause snow. They simply have a common cause - cold weather. And finally, if you keep your home at a constant temperature, the setting on your heater does not correlate with the temperature inside your home, because the temperature is constant, while you have to change the setting up and down to counteract the more or less cold weather outside. It does correlate perfectly with the outside temperature, but I'll leave proving that putting your heater on high doesn't make the temperature outside drop by 10 degrees as an exercise.
- hessenwolf 12y agoCorrelation is due to one of the 3 C's, coincidence, confounding, or causality. I just like the 3 C's.
- pleb 12y agoThe problem is that causation is an oxymoron. The measurement problem is the same as the problem of induction. Max Planck understood this, read his quotes on matter. No amount of correlation increases the probability of one event followed by another. Therefore, all scientific analysis is unverifiable. Knowledge of the world is completely unjustified. Not to mention our immense presumption of consistency in world phenomena when we really have no basis for asserting uniformity of nature. Read Hume on causality. "Belief in the Causal Nexus is superstition" - Wittgenstein If knowledge requires certainty, and our method of determining certainty is an infinite regress of reduction, it is like a recursive function which never returns a value. Logic is inherently a comparison operator, and if logic is the only mechanism available to us, we are trapped in a purely relational analysis and it is impossible for the mind to conceive of certainty beyond a non-reasonable emotional preordained state of knowing. A feeling so powerful it is personally indubitable. The premise of a singularity is invalid since it is inconceivable and the idea that causation is at any time inferred via comparison is fundamentally incomprehensible. If the premise is unclear then the argument is invalid. If I say that an airplane functions on fairy dust and you make an argument about propulsion and lift, and you claim that my assertion is wrong because I have not properly performed a rational reduction and that a detail I do not understand could invalidate my entire theory, well guess what neither of us is technically correct to any degree since the test for invalidity applies equally to both of our assumptions. The concept of proof is also an oxymoron. Hence, the very concept of phenomenal causation is an oxymoron. The notion of knowledge about the world is a disguise for a coping mechanism based on hope and desire.
- pleb 12y agoI couldn't edit, so I'll reply, but this mind bending quote really sums up the limitations of logic and observation. "Language disguises the thought; so that from the external form of the clothes one cannot infer the form of the thought they clothe, because the external form of the clothes is constructed with quite another object than to let the form of the body be recognized." - Wittgenstein
- mattfenwick 12y agoI do agree with you that science doesn't show causation, but I think your interpretation of science is incorrect: > Therefore, all scientific analysis is unverifiable. I disagree. It's verified by experiment. (Here I'm using "verify" to mean that contradictions have not (yet) been found by experimental observation). > Knowledge of the world is completely unjustified. I disagree. Again, it's justified by experiment. If that's not enough for you, too bad, because that's all we have. > Our immense presumption of consistency in world phenomena when we really have no basis for asserting uniformity of nature. I disagree. This is only asserted insofar as it's 1) useful, and 2) justified by observation. I don't think you understand the significance of scientific results. They don't say "at a fundamental, irreducible level, this is how <some system> operates." Rather, they're saying something more along the lines of "as far as we can tell, this is a description of how <some system> operates, but it's entirely possible that we're wrong. However, we don't have any data that shows that we're wrong (yet), and because this description is useful for understanding <some system>, we're going to use it." We don't need to know how a system works at a fundamental level, if understanding it at a less-detailed level is useful. I do agree with you that we can't determine "actual causality", but we don't need to, if we can instead find a bunch of really good correlations. The problem is that people find lots of crappy correlations, and/or can't tell the difference between good and crappy correlations.
- snowwrestler 12y agoPerceived correlation is usually not actual causation. That should help explain why not.
- tmbb 12y agoHey Amarok, Image Occlusion addon author here. I'd like to meet you in person (can't reply to the comment in which you mentioned me). My email is in my profile.