5 ms·
I look at these things more like "check out our classification scheme which is broken down into 4 categories", as opposed to "there are 4 different kinds of peo
by jordanlev 10y ago
I look at these things more like "check out our classification scheme which is broken down into 4 categories", as opposed to "there are 4 different kinds of people" (which the headline kinda sorta implies).
- ASpring 10y agoIf your sample is representative and your classification is supported by your sample, what is the real difference?
- ianai 10y agoYou can't use the same data to come up with and test a hypothesis.
- saulrh 10y agoDepends on what your methods are for finding your classifications. For example, if all you're doing is recording a bunch of data over some feature set and then throwing it at an unsupervised clustering algorithm like DBSCAN or single-linkage clustering, It's not unlikely that you can say that your data can be partitioned into N distinct clusters. Granted, that depends heavily on your features and your data set. It's just as easy to end up with one big undifferentiated blob, even after running PCA.
- StillHocus 10y agoI mean, that's still "who cares"? It just sounds fancier. "It's not unlikely that you can say that your data can be partitioned into N distinct clusters." Of course not...that's tautological. You can say it because that's what it means. DBSCAN isn't some wizardry that peers into the nature of the universe...its just a clump finder. Whether those clumps have any connection to anything interesting is a coincidence. You could roll a six sided die six times and run dbscan on the results and be like "holy shit! I've discovered how to partition the natural numbers 1 2 3 4 5 and 6!!"
- nerdponx 10y agoThis is an unfair and anti-scientific characterization of descriptive research. Correlation does not imply causation, but you (generally) can't have causation without correlation. Finding natural patterns of associations in data is our only reasonable starting point for finding patterns of causation. Also, developing reliable and well-researched ontologies can help other researchers when building models, making sense of other data, etc.
- StillHocus 10y agoIt's not anti-scientific at all. Looking for correlations to uncover causations in data is important, and it can certainly be the beginning in a scientific effort to build a model of the world (especially, I would argue but I guess this gets rather philosophical, when it's motivated by a belief in an underlying mechanic that drives the correlated observations, perhaps developed so fully as to be considered a 'hypothesis'). Finding natural patterns of associations in data is a totally reasonable starting point for finding a causation _when a causation exists_. Feeding arbitrary sequences of samples into dbscan and deciding that because dbscan produces output, there is causation (or, that there is an underlying phenomena that can be captured in some type of model), is ridiculous. And there are tons and tons and tons of natural phenomena that will be happy to produce clumpable inputs all the time, with no underlying behavior (including noise). I'm sure there are also tons of interesting models of human behavior that you can make out of some set of observations of human behavior via clustering. But just because you feed some data to an unsupervised learning algorithm and it discovers features doesn't imply that those features have any useful descriptive power to help us make sense of the natural world. THAT's anti-scientific thinking.
- Singletoned 10y ago> Finding natural patterns of associations in data is a totally reasonable starting point for finding a causation _when a causation exists_. So if a causation turns out not to exist then finding natural patterns was an unreasonable starting point? What would have been a reasonable starting point in that circumstance?