5 ms·
> In the first sign of something amiss, the 13,488 drivers in the study reported equally distributed levels of driving over the period of time covered in the st
by peterthehacker 5y ago
> In the first sign of something amiss, the 13,488 drivers in the study reported equally distributed levels of driving over the period of time covered in the study. In other words, just as many people racked up 500 miles as those who drove 10,000 miles as 40,000-milers. Also, not a single one went over 50,000.
I have a hard time believing that Dan Ariely didn’t know about this. The uniform distribution of mileage makes no sense, so this should’ve been caught right away. Plotting a histogram of the mileage data would’ve been one of the first things Ariely’s team did with this data.
- wjnc 5y agoIt’s not damning per se. That depends on the sampling (or otherwise it would be a tiny, tiny insurer). I’ve done studies where I sampled an insurance population to get equal group size on a few key parameters because we then did follow up questionnaires and I needed to account for non response. No point in random sampling then because all I probably would get was data on the largest groups in the population. As it turned out people loved the subject of our questionnaire (effect of preventive measures by home owners on incidence of a whole range of common claims) and we got about a 70% response rate (that’s crazy high for cold questionnaires to customers) so the study ended up quite overpowered. Not knowing the sampling, not documenting, not having the emails or at least a zip containing the work (over 4 authors)… that’s a different ballpark.
- whimsicalism 5y agoSorry, but you seem to be suggesting that non-response bias would be responsible for the uniform distribution seen in the Update mileage digits as well as the uniform miles driven distribution?
- wjnc 5y agoTheoretically, yes. I’ve used stratified sampling in this kind of research, since the underlying portfolio is so skewed.
- whimsicalism 5y agoI really don't see how any random sampling mechanism (unless you're literally stratifying based on the last digit of the odometer) would cause these sorts of results, please explain further.
- jldl805 5y agoOr, even better share the data you reference from "this kind of research".
- wjnc 5y agoThese are customers they already know I presume. So they have an inkling of the expected mileage (at least: we price on expected mileage). So you repeatedly sample dropping some samples to get enough filling in all mileage baskets. It’s a stratified sample and random. Say you have 10k customers with 1k women. You want to take a sample And want enough power to get answers on women. In that case you do stratified sampling.
- peterthehacker 5y agoThe sample size is in the article: > Nearly 13,500 drivers were randomly sent one of two policy review forms to sign… The distribution referenced is of mileage, which you’d expect to have some kind of right-skewed, continuous distribution.
- wjnc 5y agoLook I’m not defending Ariely, just saying that random sampling can be more complex than each record exactly the same sampling chance. And if the population has a few overpopulated groups but you’d like results for all groups, you don’t throw extra samples at it but use smarter sampling.