4 ms·
It's training data. There's 17 suicidal and 17 non-suicidal scans, for a total of 34 scans. They trained 34 models, leaving one scan out each time. Of those 34
by jaibot 9y ago
It's training data. There's 17 suicidal and 17 non-suicidal scans, for a total of 34 scans. They trained 34 models, leaving one scan out each time. Of those 34 models, 31 correctly predicted the left-out scan.
IANAStatistician, but this seems like a trash result.
- bduerst 9y agoHow so? Isn't there a 50% chance of getting it right by pure chance, but they got it right 91% of the time instead?
- nonbel 9y agoCross validation is ok if you do it once, but they repeatedly did it and chose the features based on the results. You can't keep adjusting your model/features based on cross validation performance without overfitting to the training data.
- jjoonathan 9y agoHow did they adjust the model/features based on CV performance? It looks to me like they did LOOCV.
- nonbel 9y agoRead the second paragraph I quoted above: "The features used by the classifier to characterize a participant consisted of a vector of activation levels for several (discriminating) concepts in a set of (discriminating) brain locations. To determine how many and which concepts were most discriminating between ideators and controls, a reiterative procedure analogous to stepwise regression was used, first finding the single most discriminating concept and then the second most discriminating concept, reiterating until the next step reduced the accuracy. A similar procedure was used to determine the most discriminating locations (clusters)." The features were chosen using the same data as used to assess predictive skill.
- yorwba 9y agoThat quote does not support your summary, unless you are basing it on the information not explicitly mentioned. (I.e. they didn't say that they were only using training data to select features, but if they are any competent, they did.)
- nonbel 9y agoSee the last part of this post: https://news.ycombinator.com/item?id=15598117 https://news.ycombinator.com/item?id=15598117 Can you provide pseudocode consistent with what they described (in the post you responding to) that wouldn't lead to leakage? I can't see it.
- yorwba 9y agoSelect a training set, leaving out one sample for validation. For all features, train a classifier on the training set using that feature. Keep the one that gives the highest discrimination score on the training set. Repeat with more features. Then evaluate the final classifier on the validation sample, which has so far not been seen in any of the steps. The result provides an estimate of the risk on unseen data from the same distribution. To get the estimation variance down, you can repeat this for all possible choices of validation sample. That means, you start the feature selection process on the new training set over from scratch and obtain another risk estimate. If they kept the features selected earlier, that estimate would be "contaminated" and not independent, but if they correctly start over, the procedure is valid.
- nonbel 9y agoMy understanding is you are saying create N (N=34 in this case) different parallel models that use different features/etc. Then take the average (or whatever summary stat) of the accuracies to get the predictive skill. When we want to use these models, we run new/test data through all N=34 models in parallel and calculate a prediction from each. Then somehow these predictions need to be combined (one again an average, etc). This is the average of the predictions, not accuracies/whatever. Where was the step combining these predictions present during the training? It seems your scheme necessarily calculates an accuracy based on a different process than needs to be applied to new data.
- wakkaflokka 9y agoIn this case nested cross-validation would have been the proper way to do this. Run your entire model selection process (scaling - feature selection w/ CV - model selection - hyper paramter tuning w/ CV) on each of the folds in the outter CV loop. That will tell you how good your process is at building a model that generalizes.