4 ms·
This statement is false, as recently demonstrated by DeepMind on retinal scans. Not only did they generalize outside of the training dataset but they were able
by bodono 7y ago
This statement is false, as recently demonstrated by DeepMind on retinal scans. Not only did they generalize outside of the training dataset but they were able to use the features learned by the model on an entirely different type of scanning device.
https://www.nature.com/articles/s41591-018-0107-6.epdf?author_access_token=PAbvHEuv_YYmrPVbG5HqKdRgN0jAjWel9jnR3ZoTv0P43NEH20hFuvBoJk6cvICihn8kmL6tmejFlnuPlbT_0KmJgK6N07SPh_ZLy0Nxb0-LAGIDBaH1fjJTkD9ahUEQpRlEudtlG9E1v3ca9xNQcQ%3D%3D https://www.nature.com/articles/s41591-018-0107-6.epdf?autho...
"Moreover, we demonstrate that the tissue segmentations produced by our architecture act as a device-independent representation; referral accuracy is maintained when using tissue segmentations from a different type of device."
- YeGoblynQueenne 7y agoIn the paper you link to, the researchers trained an image classifier on data collected from 32 sites of the Moorfiel NHS trust. The trained model was tested on, presumably held-out, data from the same dataset. This is an example of scaling a model beyond a dataset collected from a single site. It is not contrary to what I say in my comment. The researchers further tested their model on data obtained from a different device than it was originally trained on. This data was collected from the same hospital sites. The original model performed poorly on this new data and was re-trained to improve its performance. This does not demonstrate an ability to generalise to unseen data- only an ability to adjust a model to new data, by re-training.
- bodono 7y agoIt contradicts your statement: "it's not as simple as having someone at Hospital X download a pretrained model in Tensorflow and train its last few layers on some CT scans" Because in this case it was as easy as taking a model from a totally different modality and retraining the first (in this case) few layers to accommodate the new device. Furthermore the original training used 15k scans and the retraining only required 152 scans. This is totally reasonable and clear evidence of transfer and generalization. Moreover, even human operators require retraining on new devices!
- YeGoblynQueenne 7y agoMy Tensorflow comment was a bit unclear. I meant that you can't just download a generic model like the kind that is readily available, e.g. one trained on ImageNet or CIFAR etc, and expect that you can retrain it easily and get a diagnostic tool that is competitive with an expert. The models in the paper you link were specifically trained on medical imaging data. My point is that you need a lot of work to make this work even for one hospital, let alone scale to many, even more so scale at the level of a national health service. I don't see that the paper you link contradicts this. Edit: if I may summarise: I said "it's not simple" not "you can't do it". Transfer learning is not generalisation to unseen data. If the pre-trained model and the end model don't have any common instances it doesn't work [Edit: "don't have any instances with a common feature space" is more clear]. Also, you're talking about generalisation to new devices. My understanding is that this is only one aspect of the difficulties with scaling image recognition for medical diagnoses to data from different sites.