3 ms·
Thanks for weighing in, but you are completely misrepresenting my argument by reducing "Big Data" to "data that reaches a certain n" and responding to claims I
by pyduan 13y ago
Thanks for weighing in, but you are completely misrepresenting my argument by reducing "Big Data" to "data that reaches a certain n" and responding to claims I did not make.
As you are probably aware, doing these kind of analyses require more than pure statistics; they also require solid understanding of good experiment design, and this is precisely what I was arguing is at higher risk of breaking down in these types of analyses.
To clarify my previous post, what I was referring to (and what I believe is what is commonly referred to) when I was talking about Big Data is a specific albeit vaguely defined trend of analysis that tends to focus on:
a) mining data out of large, unstructured existing datasets
b) leveraging data that has been passively generated, ie. are byproducts of normal activity and not a conscious experiment design decision
c) maximizing predictive power as opposed to validating a theory
And yes, there are many ways to do bad analyses on small data sets, but that is not the point I (or, I believe, Harford) was making. The point is these types of analyses, because of their nature, tend to require additional care regarding external validity, because:
a) many make the mistake to think that big n means you don't need to worry about sampling biases, ie. what Harford was referring to as the "N = all" fallacy and what I believe was his main point; you'll agree it tends to be more common in "big n" analyses
b) since data collection is not an issue, the real challenge is in data cleaning, which requires special care because you need to think about potential biases in the way the data was generated (a process you had no control about); this can be trickier than it sounds when exploring large datasets that were passively generated, because every feature is potentially subject to these biases and some are less than obvious. The Boston case was a good example, but now consider that in many real-world datasets almost all features are subject to similar considerations (and may all be subject to different biases)
c) the focus on predictive power when using theory-free metrics leads to a risk of overfitting when the possible sources of heterogeneity are not understood (ie. the assumptions are not made explicit)
d) since they've been optimized for predictive power, they give a false sense of security ("it worked on the validation set!"); this is compounded by the fact these models will often work for a while (as is the case in GFT) before breaking down [1]
e) since there is a stronger focus on exploratory analysis, addressing the multiple comparisons problems is not as trivial as you make it sound; the issue is not the statistical tools we have at our disposal [2] but making all your assumptions explicit, which is trickier in the exploratory phase (because by doing this initial phase, you are already implicitly dismissing or selecting relationships to study)
f) since these analyses tend to be very application-oriented and to function at a large scale, mistakes have the potential to be much more destructive (for example, false positives in the Target example). This is compounded by the fact that due to the technical challenges in handling complicated data, and because applications are often found in tech companies, many people who do these analyses come from a computer science background and are not necessarily well trained in statistics or econometrics
Again, none of these are insurmountable; no one is actually dismissing Big Data analyses as a whole, but they present some unique opportunities for screwing up.
[1] Incidentally, this is precisely why I said earlier that discussing the details of GFT seemed only tangential to the point: yes, the Google researchers were well aware of the limitations of the method, and so is Harford ("Google Flu Trends will bounce back, recalibrated with fresh data – and rightly so"). The relevant point is not whether the Google researchers were right, but the false sense of certainty it instills for the consumers of the research, something I also made explicit at the end of my last post.
[2] Although some make a pretty good case that it is, but this is beyond the scope of this comment:
http://www.nature.com/news/scientific-method-statistical-errors-1.14700 http://www.nature.com/news/scientific-method-statistical-err...
Edit: I forgot to address the first part of your reply. While of course the researchers knew that people looked for these terms because they are concerned by their health, what is missing is why these specific queries are important: you may well find that some of these queries are more related to general concern, while some are specifically about treatment options, and others about vaccination. These may not evolve at the same time and in the same direction, and while they may have been indistinguishable in the past, it's entirely possible the first type of queries will be disproportionately affected by changes in the Google algorithm vs. others, or that some of the assumptions are only valid for one type of query and not the others. In economics (which is Harford's background), this is often considered insufficient when deciding whether to add a variable to a model.
- mturmon 13y agoThanks for this respectful, detailed, and analytical contribution. I agree that "Big Data" is a tendency that is worth talking about as if it is a new thing. There are edge cases that reside on the border between conventional moderate-n statistical analysis, and large-n "vacuum up lots of data and try to extract correlative information" approaches. Examples like the pothole-location collection and GFT are emblematic of something new, that's well beyond this fuzzy boundary. So it's not surprising that there are new issues. Some of the fixes may be old-fashioned, but some may not. We should also face the fact that the people doing this work often don't have any formal statistical training, so it's on the community to highlight the important pitfalls.