3 ms·
Don't you have to be concerned about differences between the structure of individual variability as opposed to population variability? Perhaps for the distinct
by mrow84 9y ago
Don't you have to be concerned about differences between the structure of individual variability as opposed to population variability?
Perhaps for the distinctions you are trying to draw it is not an issue?
- CuriouslyC 9y agoI haven't found that to be a big issue. I choose clustering and dimensionality reduction hyperparameters to penalize clusterings with few repeated individuals (high entropy), while trying to maximize the entropy of the distribution of cluster labels. To prevent cluster count explosion, you can put a regularization term/prior on the total number of clusters. An exponential or gamma with the expectation set to the square root of the number of distinct individuals is a good starting point. Of course a big part of any analysis is responding to the data, no single approach will work every time.
- mrow84 9y agoAll sounds very reasonable - thanks for the insight!