4 ms·
> 10,000 images collected from leading clinical centers internationally Lets say the data was collected from 5 different clinical centers. One risk is basicall
by argonaut 7y ago
> 10,000 images collected from leading clinical centers internationally
Lets say the data was collected from 5 different clinical centers. One risk is basically that when you deploy the model, it only works at those clinical centers due to idiosyncrasies specific to those centers. Or suppose certain doctors were more likely to take melanoma images, and certain doctors were more likely to take non-melanoma images, and both sets of doctors used different techniques to take images. These are just some ideas, there could be any number of confounding factors.
Basically, one should be default suspicious of most research papers working with small datasets (the test set is only 100 images - this is very small) that have not been deployed in the real world or otherwise validated independently. The root comment of this thread is basically saying this (a different comment is the one calling the researchers idiots).
- feral 7y ago> These are just some ideas, there could be any number of confounding factors. There _could_ be. But when the source dataset was carefully gathered for a competition and the academics are saying "Broad and international participation in image contribution ensures that the dataset contains a representative clinically relevant sample", talking about multiple different equipment and labs, things look pretty promising. A lot of ML systems are built with much less rigourous datasets and do a good job when you put them in production. 10k such images were gathered from this process. The authors then randomly select 100 images from it to use as a test dataset. 100 is a small number. But that smallness is not relevant to the _selection_ issues here. Its only relevant to whether the measurements of performance are statistically significant (i.e. that we didn't end up with a sample that by chance is particularly favorable to the ML approach.) 100 data points is enough that that's unlikely (though one should check.) Additionally the authors talk about using a much larger validation set, and performing multiple runs and checking the validation accuracy is similar. Unless they deliberately left out the damning fact that their accuracy was a lot _higher_ on their small test set than on their validation set, then its even more unlikely that the test set is a sample that is particularly favorable to the ML approach. You could argue it happens to be particularly _unfavorable_ to the humans, but that seems a stretch. Perhaps they should have created multiple test sets, and a couple of different batches of human raters etc. But honestly, their setup seems pretty good to me. > (a different comment is the one calling the researchers idiots). That's fair.
- argonaut 7y agoTheir validation set consisted of 210 positive images. The test set consisted of 20 positive images. These are very small evaluation sets for deep learning. My point is the work is promising but should be viewed with healthy skepticism (by default). I would really not read anything in particular into "Broad and international participation... ...sample." That's just a claim in a paper, it's not "the truth".
- feral 7y ago> These are very small evaluation sets for deep learning. Evaluation is a statistics question, and it doesn't matter that the deep learning model used is high capacity and needs a lot of training data. There's nothing inherently wrong with validating a complex model on a small amount of data. The paper has a section 4.2 that gives a statistical analysis. Granted, it'd be nicer if they had enough data to show statistically significant differences.