3 ms·
I get concerned about usages of big data that concentrate on "'discovery' of reliable relationships in the data." The problem being that as we add more data, t
by rscale 14y ago
I get concerned about usages of big data that concentrate on "'discovery' of reliable relationships in the data."
The problem being that as we add more data, the chances of spurious relationships increase dramatically, and the human brain is incredibly good at finding a causal explanation for those relationships, even if none exists. This can quickly turn Big Data into a noise-generating rabbit-hole, leading us down blind alleys, and wasting our time.
I love that it keeps getting easier to test our hypotheses, but a search that begins without a logical and reasoned hypothesis is a dangerous beast.
- graycat 14y agoYup. If have enough data and keep testing hypotheses long enough and keep fitting long enough, then have a good chance of finding a hypothesis can't reject and a fit that looks good, even though are looking at junk. So, divide the data in half, fit to the first half and test the fit on the second half. And if the fit fails on the second half, then what? Return to the first half, fit again, and then test on the second half again? Now want some more data to test the most recent fit. More can be done along these lines.
- pseut 14y agoI included the word reliable for a reason. That's one thing that's interesting about this stuff from a statistics perspective, how you can draw conclusions that are reliable even after some sort of search process. See, for example, the research by Joe Romano and Michael Wolf (and their coauthors) on stuff like "family-wise error rate".
- rscale 14y agoI also chose my words carefully. Some people (like you) understand that reliability and validity aren't just "p < 0.05", but that's far from universal understanding. I've seen intelligent people accept and reject hypotheses with woefully inadequate evidence, and I've also seen wild hypotheses built on the backs of strong but meaningless correlations. Dangerous beasts can be useful, but they must be treated with due care.
- nignog41 14y agoCare to elaborate how how to be more sure of reliability and validity? Any stories of inadequate examples or meaningless correlations? Just trying to learn how to better read data
- pseut 14y agoThe books by Howard Wainer, Edward Tufte, and Bill Cleveland are good starting points; it also depends a lot on your particular interests. Andrew Gelman's blog [1] is very good if you're at all interested in non-experimental data and/or poli sci applications. [1] http://andrewgelman.com/ http://andrewgelman.com/
- rscale 14y ago> Any stories of inadequate examples or meaningless correlations? A customer's marketing group was tying visitor data to geodemographic data. They put together a database with tons of variables, went searching, and found a multiple regression with a Pearson coefficient of 0.8+, a low p, decided to rewrite personas, and started devising new tactics based on the discovery. Fortunately, they briefed the CEO and the CEO said that the dimensions in question (I honestly don't remember what they were) didn't make intuitive sense, and demanded more details before supporting such a major shift in tactics. More research was done, and this time somebody remembered that this was a product where the customers aren't the users, so they need to be treated separately. And it turned out the original analysis (done without fancy analytics) was very close to correct. If the CEO hadn't been engaged during that meeting, they would've thrown away good tactics on a simple mistake. The regression was "reliable" by most statistical measures, but it was noise. A similar example holds for validity, where I saw a team make wonderfully accurate promotion response models, but they only measured to the first "conversion" instead of measuring LTV. And after several months of the new campaign, it turned out that the new customers had much higher churn, so they weren't nearly as valuable as the original customers. > Care to elaborate how how to be more sure of reliability and validity? I'm not a statistician or an actuary. I'm a guy who took four stat classes during undergrad. I know just enough to know that I don't know that much. Disclaimer aside: my biggest rules of thumb are to make sure that you're measuring the thing you want to measure (not a substitute), to make sure the statistical methods you're using are appropriate for the data you're collecting, and to make sure you understand the segmentation of your market.