4 ms·
> Many unwittingly used a data set that contained chest scans of children who did not have covid as their examples of what non-covid cases looked like. But as a
by new299 5y ago
> Many unwittingly used a data set that contained chest scans of children who did not have covid as their examples of what non-covid cases looked like. But as a result, the AIs learned to identify kids, not covid.
> Driggs’s group trained its own model using a data set that contained a mix of scans taken when patients were lying down and standing up. Because patients scanned while lying down were more likely to be seriously ill, the AI learned wrongly to predict serious covid risk from a person’s position.
> Errors like these seem obvious in hindsight. They can also be fixed by adjusting the models, if researchers are aware of them.
I wouldn't trust any method of fixing errors in data collection. You need to redo data collection and perform a proper case-control study. You need to select your cases and controls such that the demographics look similar.
In general, your goal is not to pick cases out against a background of the general population. It's to pick cases out from an at risk group.
And if there's any difference in procedure used for data collection between cases and controls, your data is compromised. Researchers often suggest correcting the dataset (cropping out artifacts etc.). In my opinion this doesn't really work... artifacts (for example in image data) can often cause subtile global differences in illumination which ML will pick up on but not be obvious to someone inspecting the images.
The only solution when you have such artifacts (and only one that would be acceptable to me if I was doing DD, or evaluating the research) is to redo data collection in the context of a more reasonable case-control study.
- LeetHacks 5y agoIf properly done, that approach would most likely generate a model that works really well and detects COVID for the right reasons. Put it out in the wild and it breaks down again because all radiologist would have to use the exact same method of scanning patients. Practice shows that scan parameters, position, annotations vastly differ across hospital, scanner and radiologist. A client of mine wanted me to look into this dataset and help creating a model for detecting COVID. One of the doctors they where working with wanted to be able to submit pictures they took (from their point and shoot camera) of the CT scan. Good luck putting that variation in your dataset.
- new299 5y agoIt's a first step I guess. At least if you can get the classification to work with controlled data collection you have reasonable evidence that the approach can work. If all other parameters area randomized between case and control, this is also fine. I'd guess you can also add illumination and other artifacts to the datasets to make the training robust to this. But ultimately in the extreme case (like wanting to submit images from a point and shoot camera) I suspect you'll have a hard time building a system that's robust to that... or robust enough to be used in a diagnostic context... I personally don't think I'd be comfortable working without a standard configuration used by the radiologists and a controlled protocol.
- forcry 5y ago>I wouldn't trust any method of fixing errors in data collection. You need to redo data collection and perform a proper case-control study. Climate research would like to have a word with you.
- new299 5y agoIt seems like a very different scenario as we don't have control Earths that we do case-control studies on...