9 ms·
Simpson’s Paradox (2016)
- freddex 8y agoI like the way this is written. Very clear and to the point, with a tone of "Hey, check out this cool thing".
- oneeyedpigeon 8y agoVery accessible, essentially making just one strong point with excellent examples and an easy-to-understand explanation. It does leave me with questions - doesn’t the number of trials in e.g. the kidney stone example count, as well as the relative success rate - but that can only be a good thing!
- TicklishTiger 8y agoThat is not a paradox. It's just the fact that a theory about something might not hold when you take a closer look at that something. In the articles example, the admission rates of a university seemed to indicate that there is a bias against women. Zooming in and looking at the admission rates of the individual departments seem to indicate that there is a bias against men. The article makes it sound like the first theory was wrong. And the second theory - the bias against men - is the real truth. Zooming in further might indicate the opposite again. Take two boxers. So far, one of them has won 86% of his fights and the other one has won 100%. According to the article, "The data is clear". Now we add more data: One fighter is Mike Tyson. He won 50 of his 58 fights. The other one is me. I did one fight in kindergarden and won it. But to be honest: I would not want to fight Tyson. As paradox as it sounds.
- rwilson4 8y agoIt’s a paradox because many people find it counterintuitive. It’s the mathematical statement of why correlation does not imply causation. The existence of a confounding variable correlated both with the purported cause (eg gender) and the purported effect (school admissions) can lead to reversals in observed association when grouped or broken out. Thus it is challenging to draw causal conclusions from observational data.
- sopooneo 8y agoI don't know, but at some point, aren't we just running up against the definition of "probability"?
- TicklishTiger 8y agoProbably.
- beaner 8y ago> By doing so, the article makes the exact same mistake Read further, the article talks about this
- TicklishTiger 8y agoTrue. Shame on me. Removed this line from my otherwise wonderful comment :)
- n4r9 8y agoIt is a paradox. In common usage, a paradox is an apparent absurdity which nevertheless holds up upon deeper investigation. In this case the apparent absurdity is e.g. "Treatment A is better at treating kidney stones despite performing worse in both trials". Sometimes the word paradox has a slightly different meaning. For example, Russell's paradox in mathematics is the opposite; it takes something apparently well-founded and shows that it is absurd.
- geocar 8y agoThat a pair of attributes doesn't necessarily exhibit independence within the universe at large, even if it exhibits independence within each sub-universe is a powerful observation, and it's a troubling one to anyone who has attempted to design a sales and marketing strategy, a drug trial, or frameworks to encourage social equality: To have it suggested I can say nothing less about these thousand students other than a thousand different things, just sounds so absurd, and yet here it is true. Sometimes people use the term "paradox" simply to a contradictory statement which upon investigation turns out to be true. In that way, "Simpson's Paradox" is absolutely a paradox.
- sopooneo 8y agoIn simples case at least, such as with the kidney stones, can we reduce our risk of reaching wrong conclusions by increasing our sample size of patients and randomizing which receive each treatment?
- rwilson4 8y agoYes absolutely! Random assigment along with statistical power and significance considerations does indeed allow one to draw causal conclusions. It’s the gold standard for causal inference.
- AnthonyMouse 8y ago> Yes absolutely! The problem with these cases is generally that people want to use data that didn't come from a controlled experiment to begin with. You have a nice, fat data set of all the people who have been treated for kidney stones -- you could never afford to do a controlled experiment at that scale. But because the treatments weren't randomized (and neither was anything else), the conclusions are erroneous. This has been a huge problem in social sciences, where you can't do the controlled experiment at all, even at a smaller scale, because there is no way to randomize the choices individuals make. All you can do is try to control for the divergence statistically -- but there isn't one confounder in real data, there are thousands or more, and each one you want to control for multiplies the measurement error (because the measurement error in the primary factor combines with the measurement error in the control factor).
- rwilson4 8y agoYou're right, and in some instances it is possible to draw causal conclusions from observational data. See [0] and [1] for two pretty different perspectives. But for this to work, you need a lot of data: both lots of units (e.g. people), and a lot of information about each individual unit. [0] Causality, Judea Pearl [1] Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction, Guido Imbens and Donald Rubin
- 8y ago
- knappa 8y agoThe sex-discrimination lawsuit against UC Berkley seems to be a kind of academic urban myth; the administration was apparently afraid of such a lawsuit and the study was done in response to those administrative fears.
- says_you 8y agoSome people would reliah the chance to disregard a narrative that fails to align with their ideology. An advantage is obtained with selective acknowlegement of reality. Now, how to go about the rationalization of ignoring it?
- techbio 8y agoIe. https://www.refsmmat.com/posts/2016-05-08-simpsons-paradox-berkeley.html https://www.refsmmat.com/posts/2016-05-08-simpsons-paradox-b...
- esquire_900 8y agoThis is the exact feeling I've been having for years, nicely described in an easy to understand language. At least in data science and (god forbid) behavioral psychology, you can answer any question any way you like - statistically valid - by slightly shifting the level of focus (as described here), definitions or angle of attack. The more data, the easier. Thanks for putting it in such a clear way :)
- throway88989898 8y agoNeatly phrased: Trends which appear in slices of data may disappear or reverse when the groups are combined.
- naasking 8y agoOr perhaps even more succinctly: slicing data can introduce bias.
- Matumio 8y agoThis is less accurate, because not slicing data can also lead to bias.
- naasking 8y agoExcept the original statement didn't make any claim about "not slicing", so neither does mine.
- throway88989898 8y agoNot slicing is nevertheless slicing. The trivial selection. Rush's song "If you choose not to decide you still have made a choice"
- IngoBlechschmid 8y agoAn explorable explanation of Simpson's Paradox, neatly complementing the article, is here: https://pwacker.com/simpson.html https://pwacker.com/simpson.html
- srean 8y agoSimpson's Paradox is one of the many phenomena that shows how different applied ML is from regular software engineering. Another one is feedback loops between decomposed subproblems. In ML encapsulation, shielding away of inner details often does not work. One needs to know what is happening on the other side of the abstraction boundary. This is a problem for managers and PM coning to ML from a purely software engineering background. They are used to encapsulation and decomposition serving them well and they expect the same.
- heavenlyblue 8y ago>> In ML encapsulation, shielding away of inner details often does not work. I call bs on this. It’s just that we haven’t yet invented a consistent type theory on top of ML.
- apathy 8y ago“Just” This would be like saying “it’s just that we haven’t proven P!=NP” in CS. Best of luck. Meanwhile applied people will deal with the problem by model diagnostics and sensitivity analysis as has been done for decades. I can’t wait for the next AI winter to come. So tired of this handwaving by people who don’t seem to have practical experience.
- srean 8y agoWhoah! You have quite a treasure trove in your favorites. The possibility of getting some work done vanished as soon as I found that.
- apathy 8y agoAlways happy to be a bad influence.
- heavenlyblue 8y agoYeah, but you just reduced the parent comment to ML being the same as CS. And the parent is saying the opposite: that they differ from each other. So meanwhile, speaking of applied knowledge... I believe you didn't even read what you're replying to.
- clircle 8y agoIirc, you can guard against simpson's paradox by designing/collecting balanced data
- GolDDranks 8y agoI thought the same; at least in the kidney stone story, the data wasn't balanced: treatment A was assigned a lot more "harder cases". Either the trial wasn't randomized or the data set size wasn't big enough.
- unparagoned 8y agoUnless you are God. You will never be able to even properly know what to factor in. Actually doing the experiment and analysis is exponentially harder. It's like saying well who cares about p=np, if you want to decrypt Aes without the key just make a super fast computer.
- currymj 8y agoJudea Pearl’s explanations of this in terms of causality are the only way it really makes sense, in my view. https://ftp.cs.ucla.edu/pub/stat_ser/r414.pdf https://ftp.cs.ucla.edu/pub/stat_ser/r414.pdf
- Gibbon1 8y agoUnless I'm misreading the take away is failure to appreciate graph/network theory is behind the Simpson paradox. And I think a lot of broken 20th century 'science'. Because theory was based on simplistic statistical analysis on processes with strong path dependence.
- lettergram 8y agoIdk why the 2016 needs to be in the title here. I understand for date relevant content, but this is not.
- city41 8y agoAnother reason to put the year is it helps people decide if they’ve read it before.
- lettergram 8y agoI doubt that helps. For instance, this is from 2016, had you read it before? I’m simply suggesting it because I don’t think it adds anything to the conversation. In addition, I’ve seen this being added more often lately and I worry it makes people think it’s date relevant (as I did) or that it somehow provides less value due to some time delay.
- city41 8y agoAnother reason is it’s possible the same author writes an update or new article on the same topic. The year helps disambiguate that.
- gumby 8y agoIt’s not uncommon for something clear and expository to be invalidated and putting the date in the title may cause someone who knows the domain to say “oh, this must be from before this was all invalidated” and post a useful reference as a comment. Well it’s a convention and some conventions (like this) are better applied uniformly than allowing for acidental editorialiation.
- mcguire 8y agoI'd like to say that the author has been reading The Book of Why, but it seems that he hasn't because he missed the punch line of the section on the paradox: you need a causal model to separate the two branches of the paradox. It's as easy to construct examples where the overall view is correct as it is so construct examples where the separate views are.
- jonahx 8y agoI'm unclear: what was the incorrect claim you're saying the author made?
- whatshisface 8y agoThe parent is not saying the author made an incorrect claim. They are saying that the parent did not continue their argument to arrive at a conclusion that someone else had, the conclusion that causal models are what tells you when you can combine datasets and when you can't.
- chii 8y ago> causal models are what tells you when you can combine datasets and when you can't. but then the causal model is subjective right? What if there are two different causal models, and a priori cannot be known which is the "true" one? Can the selection of the causal model be used to justify the dataset, in order to push a particular agenda?
- whatshisface 8y agoYour job when analysing data is simply to enumerate the possibilities and assign likelihoods to them if possible. If two models fit equally well, you're supposed to write them both down in the hope that someone will collect further data to distinguish between them. If you're cutting holes in your report for political reasons, that's just not doing the job. That's what pundits are paid to do, not (ideally at least) scientists. Fraud is easy to commit, and the fact that it's possible is not that hard of a philosophical issue.
- gok 8y agoThe last example of software optimization causing mean slowdown because users actually use the software is so true. Another example I've seen is better ML models causing accuracy to go down; users try harder things.
- jzl 8y agoCool article. My knowledge of statistics is really rusty, but isn't this another way approaching the topic of "Bayesian Thinking"? If you think about the scenarios in the article from the standpoint of predicting any given outcome in advance, male vs. female and hard department vs. easy department should be treated as "priors". Or to put it another way, Bayesian thinking means asking the question "What is the chance of X happening given Y?" A nice intro to the topic: https://betterexplained.com/articles/an-intuitive-and-short-explanation-of-bayes-theorem/ https://betterexplained.com/articles/an-intuitive-and-short-... Which explains why a positive test on a mammogram means you only have an 8% chance of having breast cancer: >The chance of getting a real, positive result is .008. The chance of getting any type of positive result is the chance of a true positive plus the chance of a false positive (.008 + 0.09504 = .10304). >So, our chance of cancer is .008/.10304 = 0.0776, or about 7.8%. >Interesting — a positive mammogram only means you have a 7.8% chance of cancer, rather than 80% (the supposed accuracy of the test). It might seem strange at first but it makes sense: the test gives a false positive 9.6% of the time (quite high), so there will be many false positives in a given population. For a rare disease, most of the positive test results will be wrong.
- currymj 8y agoThis is actually a case that shows the limits of Bayesian thinking. The power of probability is that it can work in two directions. You can use it to make predictions, from causes to effects, from past to future. Or you can use it to reason diagnostically, from effects to causes, like deducing what must have happened in the past to produce the current observation. Thinking probabilistically, these two cases are treated the same: they're both just conditioning on evidence, which is really elegant. The problem is that when the two cases really need to be treated differently, probability can't distinguish between them. For example, asking about the probability of hypothetical situations, or predicting the results of interventions. You need to know which variables are causes and which are effects, but this is outside the scope of probability. Simpson's paradox is something that only shows up when the variables involved have certain cause-effect structures. If you think in terms of these structures, it stops being counterintuitive.
- afthonos 8y agoThis is more about knowing what the right question to ask is, which is trickier than expected. In the classic example, the people who brought the lawsuit asked “what are the odds of getting into Berkeley if you are a woman?” However, if people don’t apply to “Berkeley” but instead to “Berkeley’s College of Engineering”, then the right question is “what are the odds of getting into Berkeley’s college of engineering if you’re a woman”. The paradox is due to the fact that we expect the answers to be the same. And all of this, of course, ignores sampling bias…
- jzl 8y agoObservation #2: the paradox is essentially describing statistical gerrymandering. :)
- gdne 8y agoCame here to say this. Simpson’s paradox is exactly how gerrymandering works. It’s all about how the data is grouped.
- jdhzzz 8y agoI am reminded of this XKCD comic https://xkcd.com/2080/ https://xkcd.com/2080/.
- deleted 8y ago[deleted]
- air7 8y agoThis is one of my favorite paradoxes too. Here's why: "... given the same table, one should sometimes follow the partitioned and sometimes the aggregated data, depending on the story behind the data, with each story dictating its own choice. Pearl considers this to be the real paradox behind Simpson's reversal." [0] [0]https://en.wikipedia.org/wiki/Simpson%27s_paradox https://en.wikipedia.org/wiki/Simpson%27s_paradox
- emmelaich 8y agoNot really a paradox, but you will like https://en.wikipedia.org/wiki/Anscombe%27s_quartet https://en.wikipedia.org/wiki/Anscombe%27s_quartet (if you've not encountered it before, which I suspect is unlikely!)
- YeGoblynQueenne 8y agoSo, this is the data that the wikipedia page on Simpson's Paradox cites for the Berkeley study, and that the author of the article has quoted: Men Women Department Applied Admitted Applied Admitted A [825] 62% 108 [82%] B [560] 63% 25 [68%] C 325 [37%] [593] 34% D [417] 33% 375 [35%] E 191 [28%] [393] 24% F [373] 6% 341 [7%] Above, I've bracketed in each pair of columns a) the sex with the most applicants and b) the sex with the most admissions, in a department. If that data is really the Berkeley data, then it's clear that the bias is against the sex with the most applicants, rather than either men or women. I can propose a mechanism for this kind of (with some abuse of terminology) selection bias. A department accepts some applications, then realises they've admitted too many applicants of one sex and start rejecting applicants from the dominant sex in an attempt to redress the balance. They make a mess of it and end up biased too far in the opposite direction than they originally started. Also note that in 4 out of 6 departments, more men applied than women, explaining why more departments appear biased against men (provided my observation holds). However, I can't be sure whether this is actually the original data because it's nowhere to be found on my pdf copy of the study (Sex bias in graduate admission) which I believe I got from here: https://homepage.stat.uiowa.edu/~mbognar/1030/Bickel-Berkeley.pdf https://homepage.stat.uiowa.edu/~mbognar/1030/Bickel-Berkele.... If anyone knows where this data actually comes from, I'd welcome a pointer.
- unparagoned 8y agoYou need to look at the figures. The differences that support your argument are minor and within the margin for error. You could similarly concluded that women are just smarter across the board.
- YeGoblynQueenne 8y agoI'm sorry, I don't understand your comment. What difference is minor? What is the margin for error? And how would I conclude what you say?
- _bxg1 8y agoThis is basically how gerrymandering works, isn't it?
- emmelaich 8y agoI sometimes wonder why people expect there to be any fixed, categorical semantic relationship between any set of numbers and set of natural language statements. Very rarely do the words or the numbers cover even a tiny amount of the possible interpretations.