6 ms·
I once encountered this in the real world as a data analyst a long time ago. I was working at an e-commerce company, called The Hut Group, and the whole year ou
by r_thambapillai 3y ago
I once encountered this in the real world as a data analyst a long time ago. I was working at an e-commerce company, called The Hut Group, and the whole year our marketing team had been saying our marketing cost of goods sold (the percentage of our revenue we needed to spend on marketing) had been declining across every product category. But at year end, the execs were shocked to realize that our cost of goods sold had almost doubled, from 10% to nearly 20%.
The finance team had asked me to double check the marketing team's numbers, to see if there'd been some funny math in the reporting. But the marketing team were totally right, marketing spend across the three main categories - games, beauty, and nutrition had all fallen (~15% to ~10%, ~30% to ~25%, and ~50% to ~30% respectively). However, the mix of these product categories had shifted massively, with nutrition growing from roughly 10% of our total sales to now nearly 50%.
In net that meant that whilst the marketing team had gotten more cost-efficient at selling every individual product category, the growth in the nutrition industry had vastly outstripped the growth in all other categories, and since that was the highest individual category, the aggregate marketing costs % had gone up, even though the team had improved every category. I then had the fun job of explaining the Yule Simpson paradox to a bunch of accountants.
- throwaway98797 3y agoit’s shocking that product mix wasn’t slide on reporting but marketing selects for positivity not objectivity the facts and only the facts that support what they do
- Anon84 3y agoIt’s actually surprisingly common. You can even find it in “classical” toy datasets like Iris: https://github.com/DataForScience/Causality/blob/master/1.2%20-%20Simpsons%20Paradox.ipynb https://github.com/DataForScience/Causality/blob/master/1.2%...
- appplication 3y agoCovid vaccination rates and deaths were rather famously subject to it. E.g. some combination of stats like “most covid deaths were vaccinated individuals”, “vaccination reduces death rate”, and “population segment with lowest vaccination rates has lowest covid death rates.” were all true at the same time.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- onychomys 3y agoThose aren't examples of Simpsons even taken together, but there was a famous (by which I mean it got a lot of press, including being written up in the Times and Post when it came out) study that showed that although every subgroup in Italian demographic data had lower CFRs than their Chinese counterparts, the Chinese group had a lower CFR when taken as a whole: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8791436/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8791436/
- jldugger 3y agoPretty much every dataset I work with as an SRE is full of these paradoxes. One classic published example comes from Google: A network engineer took a trip to Indonesia or something (can't find the citation to confirm the exact tale), noticed the service was slow, and when asking around everyone said "that's how its always been." Basically the local cellular networks are slow and off island fiber connects are saturated. Back at the office they decide to attack the problem by optimizing payload sizes. Does the work, reducing download sizes by half, and ships it. Latency metrics? Average and p95 latency actually increased after shipping the work to production. How does an objectively good change make things worse? Well, the service had improved for those customers so much that they used it a lot more. Even with the lighter demand on bandwidth the network latency to the datacenter was worse than typical US customers, so as more of these people realized the service sucked way less, they used it more and drove the numbers up. I have tons of these examples where a data team looks at a particular slice of request telemetry, and comes to a wrong conclusion because they didn't model enough of the system, or controlled for the wrong (or too many) variables. The worst ones the cyclic finger pointing situations that Simpson's paradox can produce: App developers blaming a regression on the server side component while the server team blames the app team, often because the server and app release schedules accidentally aligned too well. In this case we have canary data to exonerate our side of the equation, but sometimes the problem lies in even deeper spaces, like app updates from an entirely different app.
- Tomte 3y agoBut your example isn‘t a case of Simpson‘s Paradox (which is purely statistical), but Jevons Paradox (which is about human behaviour and economics).
- konstantinua00 3y agoevery time I hear about examples of simpson in peactice, I don't get what lesson to learn marketting team overoptimized, so non-nutrition demand fell? drop nutrition from line of products, so that you're both efficient in products you do and overall? these metrics are insufficient and it's better to look at gross change rather than ratios? I have no idea
- thih9 3y agoI think the last one is closest. I’d go with: “finance team should look at the gross change”, if that’s what matters for them.
- brabel 3y agoThe article suggests an answer to your question, see the last sentence of the introduction: "its lesson "isn't really to tell us which viewpoint to take but to insist that we keep both the parts and the whole in mind at once." In the case above, they failed at "keeping the parts in mind" as clearly, the different ratios between different products was crucial.
- infogulch 3y agoMaybe the lesson is to analyze different business units (product categories?) independently first, then the whole.
- gen220 3y agoIME, the "problem" (to the extent there is one) is almost always that the naïvely-chosen KPI metric wasn't specific enough. Here's a recent example from a friend. You're a SaaS company, and your home page's load time is reported as slow. You set your KPI for the quarter to be "reduce p99 load time of the home page by 50%". The load time is a function of customer size, so bigger customers = slower home page. It's actually a quadratic function. So the p99 of small customers is like the p50 of large customers. You have 20 small customers and 20 big customers. That quarter, the sales team onboards 10 new tiny customers, and 10 big customers churn. It's the holiday season in your big customers' geo, so mostly small customers are using the platform. It's the busiest time of year for the small customers, so they're over-using the platform. All these factors lead to p99 latency dropping by 60%, smashing the KPI goal. Bonuses all around, pats on the back. And no code changes needed, besides! The solution is: choose a KPI that is tightly coupled to your problem, and not confounded with other variables. In the above case, a better KPI would have been "p99 latency for large customers", because it is robust to the distribution of customer sizes across current users, churned users, and seasonal differences in usage.
- dumb1224 3y agoI thought it is pretty common to apply mixed / hierarchical linear models? I didn't study statistics but in our field of many problems of modelling biological effects we would do that. E.g https://www.pymc.io/projects/examples/en/latest/generalized_linear_models/GLM-simpsons-paradox.html https://www.pymc.io/projects/examples/en/latest/generalized_...