6 ms·
Huh? What do you mean, not working? That the AI was randomly choosing the correct race 82% of the time by luck? I'm confused by what your implying because i
by cubano 5y ago
Huh?
What do you mean, not working? That the AI was randomly choosing the correct race 82% of the time by luck?
I'm confused by what your implying because it would seem to me that the authors went through many steps to try to pinpoint how the AI was doing this identification and how baffling it was to everyone that even with a lot of x-ray information removed (8x8 pixels compared to say 4k), it somehow was still correctly picking the race.
What would this "something else entirely" that you are implying actually be?
- ceejayoz 5y ago> That the AI was randomly choosing the correct race 82% of the time by luck? No; as with the article I linked elsewhere in the thread (https://techcrunch.com/2018/12/31/this-clever-ai-hid-data-from-its-creators-to-cheat-at-its-appointed-task/ https://techcrunch.com/2018/12/31/this-clever-ai-hid-data-fr...), that the AI might have found some other indicator, like filenames in the data set, or metadata in the images that included patient name, or differences in the length of patient name (often redacted by black rectangles in x-rays in training data), or any number of other factors. This happens all the time in science. As another recent example of "whoops, turned out we were measuring the wrong thing", https://en.wikipedia.org/wiki/Faster-than-light_neutrino_anomaly https://en.wikipedia.org/wiki/Faster-than-light_neutrino_ano... Another example around AI: https://www.vox.com/recode/2019/12/12/20993665/artificial-intelligence-ai-job-screen https://www.vox.com/recode/2019/12/12/20993665/artificial-in... > One such résumé-screening tool identified being named Jared and having played lacrosse in high school as the best predictors of job performance, as Quartz reported. Are lacrosse players naturally better workers? Probably not. Are they probably whiter, wealthier, better networks, etc. than the average population? Probably. These sorts of things - as with the 8x8 pixel example - start to point to confounding variables that need to be worked out and accounted for.
- dexen 5y ago>the AI might have found some other indicator, like filenames in the data set The paper quite explicitly goes into testing and disseminating what exactly the AI detects. Two observations: - the classification clearly was primarily based on the visual content rather than spurious metadata, because various transformations of the visual content had the expected impact on classification correctness - the classification clearly wasn't based on one specific feature of the visual content but rather on multiple factors in the visuals, because various transformations to features (including masking out specific features like bone density) produced results matching expectations (usually gradual decrease in accuracy, with some thresholds). Conversely, if the classification was primarily based on factors other than the visual content, the visual transformations would have had negligible effect - possibly up to a threshold, and then would throw the AI completely off.
- ceejayoz 5y agoThe faster-than-light neutrino experiment similarly went "we've tried to account for everything we can think of and still can't figure it out" when they published. It turned out to be a measurement error. The same may be true here, and I think it's the most likely explanation. I'd be interested in whether the same model can be trained to predict patient wealth, hair color, style of clothing, religion, etc. from the same x-ray data sets.
- dexen 5y agoI am dismayed by your example that runs counter to the modern science. While "faster than light neutrino" was highly unexpected and rather suspect from the start, the "bone geometry differs slightly between ethnic groups" is well established among the anthropologists of humans. There are also parallels in wider biology of animals - mentioning that to underscore it's as scientifically expected, and not merely construed for humans alone. The question here was how exactly is AI detecting it this well from chest X-rays; the question centered around AI and possibly if it would unexpectedly influence the medical processes - rather than around the bone geometry itself. For sake of example, a random link from google search: https://www.researchgate.net/publication/24427702_Ethnic_differences_in_bone_geometry_and_strength_are_apparent_in_childhood https://www.researchgate.net/publication/24427702_Ethnic_dif...
- ceejayoz 5y agoI agree that an AI model might be able to glean race on a probabilistic basis from x-rays. This specific model's ability to do it from a 64 pixel version of said x-ray makes me skeptical it's doing so successfully.
- jcims 5y agoWhat if it was an 8x8 grayscale photo of their face? We wouldn't be particularly surprised that it can guess race. The fact that we struggle to detect patterns in the data doesn't mean they don't exist.
- shadowgovt 5y agoAs an additional comment on this point: The fact that trained neural networks cannot tell us why they give an answer and the best tool we have to explore that is to wiggle the inputs and see how the black box responds is a major concern for the whole space. Figuring out how to tag data with enough information to generate a "why" was an active area of research ten years ago and still is.
- 0-_-0 5y agoCan you explain to me how you recognize your mother's voice?
- shadowgovt 5y agoI cannot, but I don't understand why the question is asked. I'm not a convolutional neural network. And recognition of my mother's voice isn't going to have impact on, say, medical treatment for people of a particular race, or who gets a loan granted to them, or whether an autonomous vehicle successfully recognizes a person crossing the street, or whether a drone's auto targeting system decides that this blob of sensory input is a civilian sedan or a tank, or any number of other situations where the consequences of not understanding how a decision was made are that people are placing their fates in the hands of unaccountable machinery. One of the points of building these systems is to do better than human-driven.
- 0-_-0 5y agoWe rely on human judgements that can't be described all the time in many of the above situations, so why would it be a "major concern" with machines? Machines can be much more thoroughly tested than humans, so they should be more statistically predictable
- shadowgovt 5y agoThree reasons: one practical, two psychological. The practical one is that errors in a machine system scale, as do most things with machines. If I have a single bad X-ray tech who is applying the wrong medical process because I have a different race, for some reason, the damage that tech is doing is limited to whatever specific set of patients they are seeing. If a similar error occurs in a popular machine classification tool used widely by a hospital network, the damage is widespread. It is a plus that the machine can be corrected and the correction also scales, but with the (relatively speaking) stone tools we use to understand why a CNN makes its decisions these days, every fix risks breaking something else we're not testing for. The first psychological reason is that machine learning systems break in "alien" ways. They don't make the kind of mistakes humans make... They make mistakes as a product of their machinery, which means it's much much harder to predict what those mistakes will look like for an average operator. As a frequent example, it's pretty rare for humans to misclassify human beings in photographs as apes, or to fail to recognize a face in an image because the skin is too dark. That's a failure mode that happens over and over again with image recognition systems. And the second psychological reason is that humans don't trust machines to make human decisions yet. And that mistrust doesn't extend to other humans, even though we're incapable of cracking open another human's mind and understanding their thought process at the mechanistic level. It doesn't matter... we are the same organism and have a shared experience and empathy with them that we lack with machine recognition systems. It's semi-irrational, but it can't be wished away. A system for understanding why a machine makes decisions would be a step in the direction of addressing those concerns.
- FeepingCreature 5y agoI think the idea is that it's picking up on a coincidental correlational bias in the source data.
- CWuestefeld 5y agoI'm just making this up, but... Perhaps hospitals that treat a disproportionate share of poor people (which themselves are disproportionately not white), tend to use a different brand of X-ray film, and that brand has different contrast ratios than that of the brand preferred by rich hospitals. Thus, they'd be detecting the different brand of X-ray film rather than anything about the patients themselves. Of course, at this level it's still hard to imagine generating that 82% hit rate. But maybe there are multiple factors along these lines.
- lostlogin 5y ago> tend to use a different brand of X-ray film Most of us radiology folk abandoned film 20 years ago and went to digital systems (CR or DR). This doesn’t negate your query though, as vendors do have different technologies and their images do not look the same.
- kovek 5y agoThat sounds like a great idea and they can test for it! Classify the scans on the “type of film” and then alter the scan and see if the model recognizes it