43 ms·
AI Recognises Race in Medical Images
- Havoc 5y agoSurprised that’s possible given the usual refrain about it being basically just melanin
- abrichr 5y agoPrevious submission of the paper itself: https://news.ycombinator.com/item?id=28050699 https://news.ycombinator.com/item?id=28050699 We know that various features visible in medical images correlate with race, eg breast density, bone density, etc. Most likely the network is just learning a classifier on top of these features. This is trivially verifiable but conspicuously absent from the paper.
- ceejayoz 5y ago> We know that various features visible in medical images correlate with race, eg breast density, bone density, etc. That can still be picked out when pixellated down to 8x8 as illustrated in the article? That seems unlikely. I wonder if the AI is just cheating, as they sometimes inadvertently do. https://techcrunch.com/2018/12/31/this-clever-ai-hid-data-from-its-creators-to-cheat-at-its-appointed-task/ https://techcrunch.com/2018/12/31/this-clever-ai-hid-data-fr... > In some early results, the agent was doing well — suspiciously well. What tipped the team off was that, when the agent reconstructed aerial photographs from its street maps, there were lots of details that didn’t seem to be on the latter at all. For instance, skylights on a roof that were eliminated in the process of creating the street map would magically reappear when they asked the agent to do the reverse process...
- ChrisLovejoy 5y agoI suspect there must be some confounder here - like the positioning used for the CXRs correlating with race, based on the methodology used in a particular region / hospital. Seems the most likely explanation for it still working even when pixellated as 8x8?
- OneEyedRobot 5y agoThat was exactly my thinking. I wouldn't publish anything until the pixellated versions were better understood.
- SpicyLemonZest 5y agoWell, isn’t the point of publishing to get help figuring it out from other researchers in the field? I agree it’s very likely that there’s some kind of explainable trick the AI is using, but there’s no guarantee it’s an easy trick that the authors could have figured out.
- OneEyedRobot 5y agoI'd warm up to that concept if the article was: "We don't know what in the hell is going on here. Here's our source code and data set of x-rays and race. What do you think?" It could be that in the realm of machine learning, most of what is going on is people turning random knobs on a big machine and getting mysterious results. It's the birth of science without understanding.
- SpicyLemonZest 5y agoThat's precisely what the researchers are saying. In the underlying paper, they conclude that "this capability is extremely difficult to isolate or mitigate", call for "further investigation and research into the human-hidden but model-decipherable information", and suggest medical imaging people should "consider the use of deep learning models with extreme caution" until future research produces a better understanding of what's happening.
- deleted 5y ago[deleted]
- OneEyedRobot 5y agoThey always call for 'further investigation'. Looking at this: https://arxiv.org/pdf/2107.10356.pdf https://arxiv.org/pdf/2107.10356.pdf My general impression (no more than that) is a whole bunch of people crowding into a paper. The paper is mostly applying trivial image processing functions and seeing how some software they don't understand is responding. The main aim is pearl-clutching about 'bias' rather than any kind of understanding. God knows what they're going to do when any medical exam includes some kind of deep dive into the patient's genetics. No surprises. It's the nature of the era.
- abrichr 5y agoThat was my initial reaction as well but they validate on separate datasets from training which makes this unlikely. The performance despite degradation may be the same phenomenon that results in adversarial examples that are indistinguishable to human eyes, ie we know that neural nets are highly sensitive to visually imperceptible differences.
- zepto 5y agoThe article covers this.
- hgial 5y agoIt's not "conspicuously absent from the paper". They have a whole group of experiments on this: "Experiments on anatomic and phenotype confounders" and conclude "Race detection is not due to obvious anatomic and phenotype confounder variables."
- Huwyt_Nashi_070 5y agoIncredible that socioeconomic factors can even impact tissue composition and density.
- knicholes 5y agoI wonder if it has anything to do with the machines that are being used to take the images. Maybe some groups have access to one type of imaging machine where other groups have access to some other type of imaging machine.
- ttyprintk 5y agoOr some scans are ordered for a diagnostic that’s more common in one group over another. I hope that the details show that treatment and diagnostic quality are indistinguishable between race classifications.
- lostlogin 5y agoThis is a good thought, but for all of said group to have the condition would be concerning.
- umvi 5y agoI don't see why this is necessarily bad. An ML model is picking up on subtle anatomical or physiological differences between races. So what, that doesn't automatically mean the AI is racist or biased...
- ceejayoz 5y agoIt's not necessarily bad, if it's actually working. The fact that it works on an 8x8 massively pixelated version of the x-ray points to the possibility that it's not actually working, which would be bad if you based patient treatment decisions on an training set that was actually teaching the AI something else entirely.
- cubano 5y agoHuh? What do you mean, not working? That the AI was randomly choosing the correct race 82% of the time by luck? I'm confused by what your implying because it would seem to me that the authors went through many steps to try to pinpoint how the AI was doing this identification and how baffling it was to everyone that even with a lot of x-ray information removed (8x8 pixels compared to say 4k), it somehow was still correctly picking the race. What would this "something else entirely" that you are implying actually be?
- ceejayoz 5y ago> That the AI was randomly choosing the correct race 82% of the time by luck? No; as with the article I linked elsewhere in the thread (https://techcrunch.com/2018/12/31/this-clever-ai-hid-data-from-its-creators-to-cheat-at-its-appointed-task/ https://techcrunch.com/2018/12/31/this-clever-ai-hid-data-fr...), that the AI might have found some other indicator, like filenames in the data set, or metadata in the images that included patient name, or differences in the length of patient name (often redacted by black rectangles in x-rays in training data), or any number of other factors. This happens all the time in science. As another recent example of "whoops, turned out we were measuring the wrong thing", https://en.wikipedia.org/wiki/Faster-than-light_neutrino_anomaly https://en.wikipedia.org/wiki/Faster-than-light_neutrino_ano... Another example around AI: https://www.vox.com/recode/2019/12/12/20993665/artificial-intelligence-ai-job-screen https://www.vox.com/recode/2019/12/12/20993665/artificial-in... > One such résumé-screening tool identified being named Jared and having played lacrosse in high school as the best predictors of job performance, as Quartz reported. Are lacrosse players naturally better workers? Probably not. Are they probably whiter, wealthier, better networks, etc. than the average population? Probably. These sorts of things - as with the 8x8 pixel example - start to point to confounding variables that need to be worked out and accounted for.
- FourthProtocol 5y agoThere's only one race. Ethnicity may vary.
- Huwyt_Nashi_070 5y agoAny other 60-IQ takes to share?
- macksd 5y agoUnless you're the US government, then ethnicity refers to whether or not you're Hispanic, and race encompasses all other distinctions.
- dahfizz 5y agoethnicity = race race = something_else() You're just redefining words, we are all still talking about the same concepts. This is unhelpful pedantry.
- dexen 5y ago"Euphemism treadmill" [1] is a real problem when medical terminology enters common circulation. Various medical terms have shifted into insult or slur territory over the years - a risky example being "mentally retarded". Note it started as a polite expression used in stead of earlier expressions that already fell victim to euphemism treadmill - and now is understood as an insult. "Handicapped" is an example that went from polite ersatz word to bordering on insulting in our lifetimes. [1] https://en.wikipedia.org/wiki/Euphemism#Euphemism_treadmill https://en.wikipedia.org/wiki/Euphemism#Euphemism_treadmill
- goatlover 5y agoThere's one species of humans still around. Race is a non-scientific category that got made up a few centuries ago, and has continually changed over that time. Ethnicity is the regional and cultural group of usually related people that everyone prior to western colonization understood as separate groups. Often you could assimilate into a different ethnicity. So Romans would have understand themselves to be different from Greeks, Persians, Egyptians and Jews.
- throwaway894345 5y agoPlease forgive me for asking a controversial question (particularly so early in the morning), but if there are all of these biological correlations with race, what does it mean that “race is a social construct”? Is the idea that black people have greater bone mineral density (per TFA) due to social or environmental causes (e.g., diet)? For what it’s worth, I’m a staunch egalitarian and I don’t see that changing either way. EDIT: Really pleased with the largely constructive conversation in this thread. Was worried that this was going to be coopted as an ideological flame thread. Thanks for the insightful answers and good faith engagement. Keep up the good work!
- ceejayoz 5y agoIt can be both real physiological differences and a social construct, as "what level of bone density tips you from x to y" becomes a question, as does "is an x person with calcium deficiency actually a y person" sort of things that are obviously not the case.
- roflc0ptic 5y agoAs someone with a scientific bent who is of the left, I always find it incredibly frustrating when people say “x is a social construct”, because it’s technically true, but also utterly elides the dual nature of the category. Race is a social construct that can be used to infer true things (probabilities) about the real world! Other social constructs that have this property: sociology, economics, physics… This isn’t to say a lot of people who are into race science don’t wildly overstate their claims, but there isn’t literally nothing to it.
- FeepingCreature 5y agoIt's a deepity, ie. a statement with two interpretations, one true but trivial, and one enormously impactful but false.
- roflc0ptic 5y agothat is a wonderful word, thanks
- literallyaduck 5y ago"Yeah, we are going to need your chest x-rays to approve you for a loan."
- stuartbman 5y agoI'm very aware that I'm a HN novice, but can I ask why my post title was edited? The new title is much less descriptive, and x-rays are different from medical images, after all.
- andai 5y agoI don't know what the original title was, but in general the original title of the web page or paper is preferred (see the HN guidelines page). (Also, the paper covers other kinds of medical images, not just X-rays.)
- stuartbman 5y agoThat's fair enough, thanks. I'll bear that in mind for the future.
- nxpnsv 5y agoI'm struggling to understand what it's good for? Couldn't you just look? Or better yet, ask?
- ceejayoz 5y ago"Huh, why are the bacteria near that mold all dead?" "Who cares, just make another petri dish and start over." Sometimes, the utility of a piece of info is not immediately obvious. In this particular case, if there's a genuine difference between races on x-rays, it could significantly impact patient care as automated x-ray reading becomes a thing.
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- iandanforth 5y agoI asked the authors if they had compared the results with the participant's skin color. They had not. The hypothesis would be that melanin is interacting with X-rays and would explain how the system can classify "race" even at extremely degraded resolutions.
- soundnote 5y agoThis caused glorious meltdowns on Twitter. Some people just don't want to face reality.
- lostlogin 5y agoWhat reality is that? I haven’t seen the thread sorry, have you a link?
- soundnote 5y agoThat race is a physically real thing that goes beyond skin color and means pervasive differences in all parts of human biology, some meaningful, some not so much. The AI community is very invested in making a colorblind AI that can't or won't use racial characteristics in doing its job, which at least in medicine seems completely idiotic to me. We know that even at the rough proxy level of race that populations sometimes need different treatment for best results, and not building for that just leaves easy value at the table. But the ideologues building things really, really, really want it all to be totally arbitrary categories and all differences to stem from society's treatment of people. Actual biological differences are verboten.
- mlnewb2 5y agoThe only Twitter meltdowns were from your ilk. Racist Twitter exploded for some reason. It didn't seem like they (or you) even read the paper. At no point do the authors suggest making AI colorblind.
- shadowgovt 5y agoThe disquieting factor is that the network is identifying race of patients using signal that humans can't see. It's one thing if there are actual biological differences that should be factored into treatment... It's another if the AI is observing race differences that a human can't explain but as a result the AI might give different answers to other questions because it's factoring in a race signal that we can't see. We wouldn't trust "If this patient were white, I'd give this answer, but since this patient is black I'm going to give that answer... Can't tell you why, my experience just indicates to me that's correct" coming from a human doctor, so we definitely won't accept it coming from a machine. Not without a concrete reason the answer should be different if the patient is black. If I understand correctly, the paper's recommendation is for follow-up research to understand what signal the network is actually keying in to. There might be an actual biological difference. This paper wasn't able to identify it.
- andi999 5y agoAIs do not have magical abilities, I do not trust this result. AI can pick up though easily on technical artifacts. Something like a cofactor: Since they used different databases, maybe one dataset had a high number of people of one self declared race and the other the other self declared race; and each using a different intensity maximum or so.
- shadowgovt 5y agoThey accounted for specifically intensity maximum, but your overall concern is solid; I don't see anything in the paper suggesting they accounted for a full spectrum of risk factors (broad-image noise, rotation, individual "stuck pixels" that could create a hard-to-spot thumbprint in the image, for example).
- mlnewb2 5y agoSimply testing in multiple external populations already rejects this hypothesis, unless you think they all had the same scanners with the same stuck pixels. They also tested several variants of noise.
- hgial 5y agoIt might be helpful for folks to look at the blog post written by one of the authors: https://lukeoakdenrayner.wordpress.com/2021/08/02/ai-has-the-worst-superpower-medical-racism/ https://lukeoakdenrayner.wordpress.com/2021/08/02/ai-has-the... or the paper itself https://arxiv.org/pdf/2107.10356.pdf https://arxiv.org/pdf/2107.10356.pdf I see a lot of "oh it's probably just picking up on x y z" when x, y, and z are things they explicitly checked for: 1) "It's probably just the names or other metadata" – they only gave it pixel data to train on. To control for things like metadata overlaid on the image (e.g., a name written on the image) they divided the images into 3x3 sections and trained classifiers on each section separately. 2) "It's probably some artifact of how the hospital marked up the images" – they used something like 7 different datasets from different hospitals and different modalities (X-Ray and CT). If it is cheating somehow, it's not doing it in an obvious way that you can think of in a minute or two. Also note that they had more than just medical folks working on the paper; the author list includes plenty of computer scientists. It's unlikely they're making an elementary ML mistake here.
- shadowgovt 5y agoOne major risk source I see is that the size of the training data for the races isn't the same. For white vs. black patient data, there's between a 2:1 and 3:1 ratio bias in both the training and test data (and a much higher ratio bias for Asian... as high as 20:1 in some of these categories). This gives the CNN more information on one race than another, which can create a classifier that performs very well on the training and test data it has access to but then flakes spectacularly on data outside the training set (because the source isn't representative of the total variance in the global population).
- mlnewb2 5y agoThey tested on tons of different external datasets, and at least one of the training datasets was balanced. Same results were obtained.
- shadowgovt 5y agoThis is probably a great time to remind everyone that the reason the blood types are A, B, AB, and O (as opposed to, say, A, B, C, D or another nomenclature) is that when the first blood type experiments were run, only people with A and B protein configurations were available for testing in the lab where the tests were executed. I'd be very cautious drawing sweeping conclusions from research like this. The researchers have a heavy burden to prove that what they don't mean is "recognizes race in this training dataset."
- wswope 5y agoI thought this was some expert-level trolling at first, but your post history suggests you’re acting in good faith, so: Blood types are categorized as A/B/AB/O because of the presence (or for O, absence) of protein markers on the surface of blood cells. A/B/C/D would be a much less descriptive system.
- deleted 5y ago[deleted]
- shadowgovt 5y agoThis is what I get for trying to tell that story from memory. What I intended to say was that the original blood groups were A, B, and C, and only with later research into the antibodies and surface proteins was it discovered that 'C' was really an 'AB' group. 'O' wasn't in the original set because none of the original donors had type-O blood, nor was Rh-factor discovered for the same reason. The original clotting discoveries weren't wrong, but they were dangerously not-right, as in "Your blood turns to jelly in your veins" not-right. The relevant point is that early research (especially when there isn't a well-understood causality story, as is the case here) is more likely to be off due to small sample size than representative of the reality for the global population. I would treat a claim as broad as "AI Recognises Race in Medical Images" as the kind of thing that will end up with giant qualifiers on it as follow-up work is done.
- desktopninja 5y agoPrevious discussions: https://hn.algolia.com/?query=AI%20has%20the%20worst%20superpower%20medical%20racism&type=story&dateRange=all&sort=byDate&storyText=false&prefix&page=0 https://hn.algolia.com/?query=AI%20has%20the%20worst%20super... Personally think 'race' is nothing more than a fantastical vanity construct. Really its tribalism. Furthermore, I find it hard as well to comprehend how it holds weight in the medical industry. Race is not real science. Race is entertainment science. AI is mostly entertainment science.
- desktopninja 5y agoRespectfully to the downvotes, race is an unreliable identifier as mixed race populations completely upend the results: https://www.pinterest.com/Biraciality/ https://www.pinterest.com/Biraciality/ Is there an American race? What does a pure American look like? The term race is a self-identifier (what tribe do i belong in) and not biological. Really the same thing as saying "I'm a northerner" or "I'm a southerner". If AI were to sample datasets from each group, we'd unreservedly say there's a northern race and a southern race.
- JoeAltmaier 5y agoNever mind the images; who was deciding what 'race' the training data matched against? In this modern age of globalism, they must have searched hard to find anyone with any kind of historically-categorized dna. I'm guessing, they just used folks' self-identification for race on some form. Which is largely a social construct.
- drocer88 5y agoFrom the preprint ( https://arxiv.org/ftp/arxiv/papers/2107/2107.10356.pdf https://arxiv.org/ftp/arxiv/papers/2107/2107.10356.pdf ) : " In this work, we define racial identity as a social, political, and legal construct that relates to the interaction between external perceptions (i.e. “how do others see me?”) and self-identification, and specifically make use of the self-reported race of patients in all of our experiments. "
- JoeAltmaier 5y agoThank you. Exactly as I feared. Its a social definition, and makes it even more weird that an AI could predict from physical attributes i.e. an x-ray.
- HideousKojima 5y agoWhat percentage of dark skinned people of African descent do you think wouldn't self-identity as black? Same for percentage of light skinned European descent, Asian descent, Central/South American descent, etc.? I'm sure there are some out there who would identify as a different race, and there are people of mixed heritage that throw a wrench in the works too, but the vast majority of humans will self-identity as a race that matches with their skin color and heritage.
- JoeAltmaier 5y agoAnd there lies the fundamental socially-embedded racism inherent in the idea of 'black'. That somehow, having even a fraction of an ancestral African lineage makes a person 'black' - it taints them, stains them and they are automatically labelled as that part, even though they may be 95% other lineages. For clarity: Why isn't a person who has even 1/64th of a 'white' heritage automatically termed 'white'? It makes exactly as much sense that way. Yet we choose the other. It's clear that an AI cannot know about this cultural bias. So how on earth can it map an x-ray to the imaginary idea of 'black'?
- lmilcin 5y agoAnd... they found it looks at the name of the patient on the border of the image or something similar. Like the time some team tried to evolve an FPGA net to solve some problem efficiently with a genetic algorithm and it learned to use a bunch of FPGA transistors as an antenna to communicate with another part of FPGA chip through interference. Unfortunately, it would not work on other FPGA chips even from the same lot.
- lostlogin 5y ago> they found it looks at the name of the patient on the border of the image or something similar. The images were anonymised. If they AI cheated, it’s not obvious how. The paper is interesting. https://arxiv.org/abs/2107.10356 https://arxiv.org/abs/2107.10356
- askesum 5y agoGreyhound and Schaefer are separate races. The fastest Schaefer would lose a race with the slowest Greyhound. Jamaicans seem to be faster than swedes. But still, the fastest swede is faster than almost every jamaican. Swedes and jamaicans are not separate races.
- motohagiography 5y agoAny clustering similarity scheme for biometric data would yield similarity categories that we may or may not name "races" though. We could probably do the same with text analysis, where the emergent distinct flavours would create categories. A previous HN story that did specifically this (https://news.ycombinator.com/item?id=27568709 https://news.ycombinator.com/item?id=27568709) could have just as easily been called "tribes." The bigger question is whether the categories provide heuristics with valuable predictive illumination. "Valuable," being the key term to solve for. Ethnicity information in medicine may be a fast heuristic for testing for things like melanoma and diabetes, but even that this fast sorting rule might provide a time/steps shortcut or intuitive leap to test for a diagnosis is likely really more an artifact of the cost of testing and examination than the result of a physical/biological determinant. I'd conjecture that a world with tricorders where the cost of scanning for disease is equal and controlled, would likely yield results that were less-ethnically correlated - and then edge cases that were exclusively ethnically correlated, e.g. over a very polarized distribution. There's also the question of whether the tricorder measures complete things, and who decides. This is to say, there are differences and combinations that may aggregate into categories, but the meaning of the differences is dynamic, subjective, and a function of what level of abstraction you are looking at them from. E.g. at the level of a statement like "most foo people are bar," you've already cancelled out most of the information about your sample, so the coherence of something that low-information is going to be limted as well. In this sense, the "social construct," description is a response to these noisy dynamics, and it's consistent to a point. In this view, race is only ever a determinant when we let it be, as the result of chosen and learned interpretations of these cognitive grouping dynamics. When the cost of errors is low, we can afford to unlearn these abstractions. Modernity and civilization implies the cost is low. Taking that further, when the real cost of errors is high enough, you get a reinfocement effect on the bias where the surviving population is made up mainly of people who exercised that fast heuristic (hence long-lived homogenous populations), because the tolerant ones evoltionarily select out as a result of that high error cost. I could even extend this further to define racists today as people who percieve a high cost to being wrong in their generalizations, which correlates well with being poor, but also, very rich, just less so in the middle between. Anti-racism becomes a kind of signal that shows you can afford to be wrong, and oddly, racism in this model is intended to signal you have a lot to lose. If you want to reduce racism, solve for the security issues for people who percieve a high cost to being wrong about openness. If you want more racism, just antagonize people who percieve that they have a lot to lose. I'd wonder how well that generalizes.
- lostlogin 5y agoWe were marvelling at a surface shaded render made on a new Siemens MR from an T1 MPRage on a very still and compliant patient. It basically looked like a black and white photo (though with the tools at hand we could cut the image in half and look at the brain). You could see the facial hair and you could identify the patient if you knew them. Medical imaging is moving along at pace and it would be interesting to see what could be inferred from a dataset of images of this quality.
- threshold 5y agoYou want AI reviewing medical imaging to recognize race because the likelihood of certain diseases is higher for some races than others.
- poulpy123 5y agoAFAIK there is no scientific definition of race so I don't see what could be recognised by an algorithm
- mlnewb2 5y agoIn medicine there is an accepted definition which is used by, for example, the national institutes for health in the US. They use that to define "health disparities" between different racial groups. That definition is what the researchers used here. Race as a social construct, self-reported by the patient.