4 ms·
The authors refer to a literature describing "shortcuts" as "correlations that are present in the data but have no real clinical basis, for instance deep models
by carbocation 2y ago
The authors refer to a literature describing "shortcuts" as "correlations that are present in the data but have no real clinical basis, for instance deep models using the hospital as a shortcut for disease prediction". It feels like a parallel language is developing. Most of us would call such a phenomenon "overfitting" or describe some specific issue with generalization. That example is not a shortcut in any normal sense of the word unless you are providing the hospital via some extra path.
They call demographics like age and sex "shortcuts" but I find this to be a frustrating term since it seems to obscure what's happening under the hood. (They cite many papers using the same word, so I'm not blaming them for this usage.) Men are typically larger; old bones do not look like young bones. There is plenty of biology involved in what they refer to as demographic shortcuts.
I think you could take the same results and say "Models are able to distinguish men from women. For our purposes, it's important that they cannot do this. Therefore, we did XYZ on these weakly labeled public databases." But perhaps that sounds less exciting.
- bravura 2y agoThis is not overfitting precisely because of the bias/variance tradeoff. A model overfits if it is unnecessarily COMPLEX for the training data. If there is bias in the training (and validation and test) data that allows a SIMPLE model to fit the data because of a spurious correlation, that is not overfitting.
- carbocation 2y agoThere is no reason to believe that an X-ray model's correct estimation of age is due to "spurious correlation". Rather, it seems to be "undesirable correlation".
- TeMPOraL 2y ago> Rather, it seems to be "undesirable correlation". More specifically, politically undesirable correlation - as in, "it's there, but its existence upsets some people". It's pretty obvious and self-evident that there are meaningful biological differences related to age, sex, and other demographics. Whether or not they're clinically relevant for a specific diagnosis under question is one thing, but they are clinically relevant for great many diagnoses; trying to "de-bias" reality here will only lead to unnecessary suffering and loss of life.
- knallfrosch 2y agoIt seems that the models not only use bone structure (or similar) itself, but improve their prediction with "forbidden" knowledge, such as "green people have an overall lower or higher rate of bone cancer than others" or "people who come to this specialized hospital have bone cancer anyway, so I don't even look at the image" Now you can say that this is perfectly fine and represents the most likely real-world use case. Or you might prefer a model that looks at the image only, with the implicit assumption that this "forbidden knowledge" will be added by human doctors later on in the pipeline. This is beneficial because the "forbidden knowledge", such as whether patients from Hospital A always have bone cancer, might change overnight! Imagine the hospital gets assigned a new name in the system and the prediction is shit now. This second, "unbiased" AI system will always have a worse performance, because you lobotomize it when you kill the forbidden knowledge with a sledgehammer. This study just showed that "group fairness" is at odds with optimal predictions" and how much it is at odds. PS: You might even prefer a society where everyone is worse off, but every protected group is equally bad off. You'd also ban the humans from applying the forbidden knowledge. Whether that is desirable, is, of course, out of the scope of the paper.
- carbocation 2y ago> Or you might prefer a model that looks at the image only These models are only looking at the images. They are inferring demographics.
- RandomLensman 2y agoYes, these things can be a factor for a specific diagnosis but why use AI when (just) going back to conditional probabilities based on groups instead of making a true individual diagnosis? You want each diagnosis to be correct and not just a good average.
- TeMPOraL 2y agoYou always diagnose on conditional probabilities. The diagnosis is conditioned on your belief in occurrence of symptoms, which is conditioned on the observations and results of tests you make. In an ideal case, you observe well enough to make a definite diagnosis; in real case, there's always some uncertainty, plus you can't do all the tests simultaneously - which is where all those proxy factors like demographics are useful: they help prioritize tests and narrow down the diagnosis quicker.
- pbhjpbhj 2y agoDo the sorts of diseases ML is being used to detect have a flat incidence profile over age? Even if they do, negating other diagnoses that are age dependent would still mean ML models would acquire a measure of patient age, say.
- carbocation 2y agoAlmost every noncommunicable disease of adulthood becomes more common with age. (The diagnoses in this paper were things like "cardiomegaly" which are mostly just X-ray findings and, while they have ICD codes, are not a meaningful diagnosis that a practicing physician would care about.)
- ivanbakel 2y agoI think you're oversimplifying the issue. It's not important that the model cannot distinguish between people of different demographics, it's important that the model does not use demographic information in place of actual diagnosis for the sake of better accuracy. That the model can determine biological sex from X-rays wouldn't be an issue if it never shortcuts the diagnostic process by using biological sex in place of meaningful diagnostic data. I would not like a model to ignore a melanoma in my chest scan because it can deduce that I was born male and my risk of breast cancer is quite low. The idea of penalising a model which takes such biological shortcuts (because its subgroup accuracy gets worse) seems like a good solution, and it's cool that the approach works in TFA.
- carbocation 2y agoI believe that imbuing what the model does to make a prediction with the idea of "shortcuts" is oversimplifying a more complex issue. I don't think it's helpful to describe a model that can distinguish demographics as taking "shortcuts". I think that adds a layer of jargon that we then need to cut through to understand what is going on. There are plenty of tools that have been developed to enforce that models perform in a manner that is unbiased across some dimension (e.g., sex, or hospital, etc). (For example unsupervised domain adaptation.) I think that splitting the field with jargon makes it more difficult to follow the breadth of the field.