4 ms·
That author fails to grasp the concept of good enough, or optimal marginal return. Subsampling is often done because you can get the level of prediction desire
by tjpaudio 8y ago
That author fails to grasp the concept of good enough, or optimal marginal return. Subsampling is often done because you can get the level of prediction desired at a much lower cost, both in terms of hardware and employee hours. Those projects that can glean value from full data usage certainly go for it, but always using full data on everything would be a poor decision.
- privong 8y ago> Subsampling is often done because you can get the level of prediction desired at a much lower cost, both in terms of hardware and employee hours. While that's valid, I think the author's concern is that authors frequently do not demonstrate that the subsampling is representative. So the conclusions from the sampled data my not be as accurate as is claimed. (edit: minor phrasing change)
- nimish 8y agoIf only there was an entire field that exists to characterize when subsampling a population is valid and sound.
- yourMadness 8y agoWhy spend years learning those results when you can ignore them at no cost? And have a happy client that finally gets the results he wanted when the other analysts said it was impossible.
- pedasmith 8y agoIf only my non-statistical peers would recognize that sampled=fast and fast=more checks and explorations. It's like they recognize that fast compile times are a great thing (I've had to wait in line with punchcards, and it sucks), but they are completely oblivious to the exact same argument when it comes to data.
- deleted 8y ago[deleted]
- cwyers 8y agoElsewhere he talks about how subsamples are chosen because they take but seconds to run instead of minutes. If a sample is small enough that you're saving that much time, and analysis is that cheap, you can still do five-fold or ten-fold cross validation in less time than the full data set analysis and get a very good idea on if your subsample is representative of the data or not.
- triplee 8y agoThat's what I took from it. Forbes is aimed at the people buying solutions so from the perspective of a CIO or CTO (or even the people they're supporting with their analysis systems), the article is telling them that they may not be getting what they think they paid for. This is about industry trends, and you CAN get a representative sample in reasonable time, for some definition of reasonable. The takeaway from this article is that what someone in the data space may think is reasonable isn't what someone who just paid for an army of data scientists and a data lake solution because those things are sexy thinks is reasonable.
- pixl97 8y ago"We randomly sampled swans in the data set, they are all white" "Then explain this black swan"
- jaclaz 8y ago>"Then explain this black swan" It's not black, it is a very dark white, and you might be looking at it with the wrong lighting.
- PLenz 8y agoIf you are doing things right you never say the equivalent of "they're all white" - you give a distribution. Explaining that and what it means is a communications issue, not a data one.
- geebee 8y agoYou do raise a fair objection, but "black swan" events are a known issue where it comes to sampling. For example, an answer to your question could be: "you asked us to find the average wing span to beak size ratio for male and female swans. Including the black swan doesn't change it at all, and the massively larger sample set doesn't improve our accuracy".
- dredmorbius 8y agoI didn't get that sense. I think he 1) does understand the value of sampling but 2) is highlighting the false advertising that's often practiced of claiming Big Data based results when what's actually presented is in fact sampling.