5 ms·
>the AI might have found some other indicator, like filenames in the data set The paper quite explicitly goes into testing and disseminating what exactly the A
by dexen 5y ago
>the AI might have found some other indicator, like filenames in the data set
The paper quite explicitly goes into testing and disseminating what exactly the AI detects. Two observations:
- the classification clearly was primarily based on the visual content rather than spurious metadata, because various transformations of the visual content had the expected impact on classification correctness
- the classification clearly wasn't based on one specific feature of the visual content but rather on multiple factors in the visuals, because various transformations to features (including masking out specific features like bone density) produced results matching expectations (usually gradual decrease in accuracy, with some thresholds).
Conversely, if the classification was primarily based on factors other than the visual content, the visual transformations would have had negligible effect - possibly up to a threshold, and then would throw the AI completely off.
- ceejayoz 5y agoThe faster-than-light neutrino experiment similarly went "we've tried to account for everything we can think of and still can't figure it out" when they published. It turned out to be a measurement error. The same may be true here, and I think it's the most likely explanation. I'd be interested in whether the same model can be trained to predict patient wealth, hair color, style of clothing, religion, etc. from the same x-ray data sets.
- dexen 5y agoI am dismayed by your example that runs counter to the modern science. While "faster than light neutrino" was highly unexpected and rather suspect from the start, the "bone geometry differs slightly between ethnic groups" is well established among the anthropologists of humans. There are also parallels in wider biology of animals - mentioning that to underscore it's as scientifically expected, and not merely construed for humans alone. The question here was how exactly is AI detecting it this well from chest X-rays; the question centered around AI and possibly if it would unexpectedly influence the medical processes - rather than around the bone geometry itself. For sake of example, a random link from google search: https://www.researchgate.net/publication/24427702_Ethnic_differences_in_bone_geometry_and_strength_are_apparent_in_childhood https://www.researchgate.net/publication/24427702_Ethnic_dif...
- ceejayoz 5y agoI agree that an AI model might be able to glean race on a probabilistic basis from x-rays. This specific model's ability to do it from a 64 pixel version of said x-ray makes me skeptical it's doing so successfully.
- jcims 5y agoWhat if it was an 8x8 grayscale photo of their face? We wouldn't be particularly surprised that it can guess race. The fact that we struggle to detect patterns in the data doesn't mean they don't exist.
- ceejayoz 5y ago> We wouldn't be particularly surprised that it can guess race. That's actually a great example of this problem, though. https://www.theverge.com/21298762/face-depixelizer-ai-machine-learning-tool-pulse-stylegan-obama-bias https://www.theverge.com/21298762/face-depixelizer-ai-machin... > It’s a startling image that illustrates the deep-rooted biases of AI research. Input a low-resolution picture of Barack Obama, the first black president of the United States, into an algorithm designed to generate depixelated faces, and the output is a white man. > It’s not just Obama, either. Get the same algorithm to generate high-resolution images of actress Lucy Liu or congresswoman Alexandria Ocasio-Cortez from low-resolution inputs, and the resulting faces look distinctly white. As one popular tweet quoting the Obama example put it: “This image speaks volumes about the dangers of bias in AI.”
- Jensson 5y agoWanting the AI to have the same racial bias as Americans, that is that Barrack Obama who is half white half east African should be categorised the same as someone with west African heritage is just dumb. Barrack Obama has a white mother and a father from Kenya, so has little resemblance to African Americans who have mostly west African heritage, only reason people don't see a difference is because they are so used to categorise people by skin color. Of course getting data from all areas of Africa and all mixes of people would be great, but there are limits and adding in more west Africans wouldn't have helped accurately depixel Obamas face.
- hgial 5y agoForgive me if I don't consider "you can tell someone's race from physical features" to be quite as extraordinary a result as "particles can travel faster than light".
- ceejayoz 5y agoThe point is less "they're equally significant conclusions" and more "sometimes the thing you thought you discovered isn't a thing".
- vilhelm_s 5y agoIn principle yes, but did you read the paper? They do a lot of completely crazy things like blurring the image until it's just fuzzy blobs, or doing a high-pass filter on it until it just looks like noise (they comment that a human could not even guess that it's an x-ray picture), and they still get very high accuracy. Basically no matter what they try they can still get the race out, with slightly lower percentage numbers. When reading it I also thought this is too good to be true, and they may have some kind of bug in their code...
- dexen 5y ago>crazy things like blurring the image until it's just fuzzy blobs That... that doesn't influence one of the presumed ways the NN categorizes images: the trend in bone geometry. The "blobs", while fuzzy, still largely retain the relative proportions to each other. Or, in other words, proportions of image elements are invariant for operations of scaling and of blurring.
- andi999 5y agoIf that is the case you could write a normal algorithm to get the proportions and see if it separates you data set nicely. (which should be done to prove this assertion)
- drdeca 5y agoIt seems clear that if you just, instead of blurring the image, set the images (or the part of the image with the x-ray scan) to the same image, then that would work to evaluate whether it is getting the information from the image or from some other source. Seeing as this would be easy to do, I imagine that if it is at all plausible from what they know that it is getting information from anything other than the x-ray scan, that they would have already tried this? I do wonder how good of a predictor something would be if it just went off the average brightness of the image. Probably very bad, but maybe better than chance? Well, better than chance on the training set is to be expected, the question I guess is whether it would be better than chance on the test or validation set (I’m not confident in my understanding of the distinction between testing set and validation set. Is the idea that if you are using the score on the testing set to decide when to stop training, and maybe what hyper parameters to use or something, and other things to determine which model, you only try the model on the validation set once you have decided on your final version of the model?)