8 ms·
It's a hard problem to work around which is rooted in the data available. I published this paper while I was at Google: https://www.nature.com/articles/s41591-0
by dontreact 6y ago
It's a hard problem to work around which is rooted in the data available. I published this paper while I was at Google:
https://www.nature.com/articles/s41591-019-0447-x https://www.nature.com/articles/s41591-019-0447-x
The only data we were able to get at the time was mostly white patients. We talked to many hospitals but many were/are reluctant to share anonymized data for research. I'm not at Google so I'm not sure the status of the project now, but there was a real attempt to try and gather more diverse data. Unfortunately there were a lot of obstacles put up by those who have the data (hospital systems).
Fundamentally, it seems to me like there just aren't as many lung cancer screening scans out there for non-white patients as there are for white patients. How do we get around this? How do we improve on the situation? I fundamentally believe that machine learning in the long term can make medicine more accessible to more diverse groups, but not if we shoot it down out of fearmongering right away.
I agree that bias is a problem, but part of what needs to happen to get more diverse data is simply having more data available. There is real promise in this technology and if we have a one dimensional view of it ("is it or is it not dangerous because of bias/privacy") then we will fail to get past the initial humps related to fear and distrust.
- aisofteng 6y agoAs a fellow practitioner, I entirely agree. Actually, reading this article made something click for me regarding the oft discussed and denigrated “bias in AI” always brought up in discussions of the “ethics of AI”: there is no bias problem in the algorithms of AI. AI algorithms _need_ bias to work. This is the bias-variance trade off: https://en.m.wikipedia.org/wiki/Bias–variance_tradeoff https://en.m.wikipedia.org/wiki/Bias–variance_tradeoff The problem is having the _correct_ bias. If there are physiological differences in a disease between men and women and you have a good dataset, the bias in that dataset is the bias of “people with this disease”. If there is no such well-balanced dataset, what is being revealed is a pre-existing harmful bias in the medicinal field of sample bias in studies. If anything, we should be thankful that the algorithms used in AI, based on statistical theory that has carefully been developed over decades to be objective, is revealing these problems in the datasets we have been using to frame our understanding of real issues. Next up, the hard part: eliminating our dataset biases and letting statistical learning theory and friends do what they have been designed to do and can do well.
- jjcon 6y ago> AI algorithms _need_ bias to work. This is the bias-variance trade off: https://en.m.wikipedia.org/wiki/Bias–variance_tradeoff https://en.m.wikipedia.org/wiki/Bias–variance_tradeoff To be clear, statistical bias is in fact distinct from the colloquial term ‘bias’ most people use - but they can be interpreted similarly if given the proper context (which you did)
- YeGoblynQueenne 6y agoIn machine learning the "bias" that relates to the bias-variance tradeoff is inductive bias, i.e. the bias that a learning system has in selecting one generalisation over another. A good quick introduction to that concept is in the following article: Why We Need Bias in Machine Learning Algorithms https://towardsdatascience.com/why-we-need-bias-in-machine-learning-algorithms-eff0343174c0 https://towardsdatascience.com/why-we-need-bias-in-machine-l... The article is a simplified discussion of an early influential paper on the need for bias in machine learning by Tom Mitchell: The need for bias in learning generalizations http://dml.cs.byu.edu/~cgc/docs/mldm_tools/Reading/Need%20for%20Bias.pdf http://dml.cs.byu.edu/~cgc/docs/mldm_tools/Reading/Need%20fo... The "dataset bias" that you and the other poster are discussing is better described in terms of sampling error: when sampling data for a training dataset, we are sampling from an unknown real distribution and our sampling distribution has some error with respect to the real one. This error manifests as generalisation error (with respect to real-world data, rather than a held-out test set), because the learning system learns the distribution of its training sample. Unfortunately this kind of error is difficult to measure and is masked by the powerful modelling abilities of systems like deep neural networks, who are very capable at modelling their training distribution (and whose accuracy is typically measured on a held-out test set, sampled with the same error as the rest of the training sample). It is this kind of statistical error that is the subject of articles discussing "bias in machine learning". Inductive bias has nothing to do with such "dataset bias and is in fact independent from dataset bias. Rather, inductive bias is a property of the learning system (e.g. a neural net architecture). Consequently, it is not possible to "eliminate" inductive bias - machine learning is impossible without it! The two should absolutely not be confused, they are not similar in any context and should not be interpreted as in any way similar.
- epmaybe 6y agoI certainly see and empathize where you are coming from. However, I would like to add that it kind of makes sense that you’d have more white people with scans available. Focusing on the USA for a second (and note that this likely applies elsewhere too, since screening programs are really only in full force in developed countries which, surprise surprise, are predominantly white). Non white patients don’t get screened as much. Non white patients don’t go to the doctor as much. Non white patients are inherently fewer than white patients. I agree that finding a good way to get anonymized data is going to help in future endeavors, but we do need to keep in context the players involved in getting and using that data. And of course the ultimate goal, to improve health regardless of race, social class, wealth, etc.
- jsinai 6y ago> Non white patients are inherently fewer than white patients Look at global population statistics. While there are no official global figures for ethnicity, we can make some simple inferences based on continental distribution [1]: North America + Europe combined (17.19%) is barely as much as Africa (17.2%), and this is ignoring the fact that a good part of the North American population is non-white. There is nothing "inherent" about there being less non-white patients. The issue is inbalanced access to health care and screening programmes, but that is not inherent. This is without even mentioning that Asia accounts for almost 60% of the global population. [1] https://en.wikipedia.org/wiki/Demographics_of_the_world#2020_population_distribution https://en.wikipedia.org/wiki/Demographics_of_the_world#2020...
- epmaybe 6y agoI think I am agreeing with you, based on your comment on imbalanced access to healthcare and screening programs. I’m saying the same thing in that data collection for ct scans is really only happening in countries that are predominantly white, not that it isn’t possible for other countries to implement programs and collect that data for training purposes. Edit: unless of course you have found large databases that suggest my intuition is wrong?
- usrnm 6y agoHonest question: does it really matter for lung cancer? Is there much difference between races in this particular field?
- dnautics 6y agoHow would you know without the data? There are plenty of medical conditions with wildly divergent rates and pathophysiologies based on human genetics.
- JimTheMan 6y agoI don't think the burden is on us to prove a negative. It'd be way better to think of it in terms of genetic markers instead of races. Race in general practice is a social construct based on colour, appearance and culture. It's a level of abstraction away from the actual genetic data that we don't need with our level of technology. We could dig right into the genes and throw away our outdated notions. That's where the actual useful detail is. Who cares about the color of the person if they have the same amount huntington repeats you know?
- dnautics 6y ago"Race in general practice is a social construct based on colour, appearance and culture" That is a very dangerous lie. If you are a doctor I pray for your patients that you are not taking a colorblind approach because their lives could be at stake. It just so happens that the way people look correlates with the diseases they get and sometimes don't get. Sometimes you can't sequence the person's genome because it's too time consuming. Sometimes we don't even fucking know which genes are responsible. Your example of huntington's is quite frankly the outlier, and a facile example (you know it is because that's what they teach in 8th grade biology class). And even that's a shitty example because the severity only roughly correlates with number of tandem repeats. Any given person can tolerate a load of misfolded huntingtin, and their tolerance is probably governed by a raft of other conditions, that are familial and do track with race (venezuelans are for some reason more sensitive to huntingtin repeats). Other conditions, too, transthyretin amyloidosis is more common among japanese and finns and really rare among africans (except for one variant that causes congestive heart failure; but in finns and japanese it presents as liver failure), etc etc etc.
- jsinai 6y ago> Fundamentally, it seems to me like there just aren't as many lung cancer screening scans out there for non-white patients as there are for white patients. Just to qualify, you mean for the USA alone? It seems to me that part of the challenge is recognizing that the research needs to take place beyond just Western countries, or acknowledging it where such research is already occurring. Understandably many people would not be so comfortable with Google accessing patient data from around the world, so the next challenge is how diverse and global data can be protected so that important medical research can take place without any compromise of privacy. The challenge is hard but surely not impossible, as this was the approach taken by the AstraZeneca-Oxford (and others) which conducted part of its covid vaccine trials in South Africa to test efficacy on non-white populations.
- dontreact 6y agoAt the time we were conducting this research lung cancer screening existed mostly in Europe, China and the U.S. Note that if we had to conduct a 5 year multi site lung cancer screening trial ourselves in addition to doing the research, there would be basically no way of getting private funding for that. Those trials are very, very expensive and take several years to reach a conclusion. Add to that the potential optics of Google “experimenting” in developing countries and the blowback risk from that...
- DoreenMichele 6y agoHow do we improve on the situation? Given economic realities and racist history (consider what happened in Tuskegee as one example), in the US you would need to provide free screenings to poor people under circumstances that convinced people of color they can trust you while signing the documents to let you have their data. This is a fairly high bar to meet and one most studies are probably making zero effort to really meet. I'm part Cherokee and I follow a lot of Natives on Twitter due to sincere and legitimate interest in my Native heritage, but the world deems me to be a White woman so I am sometimes met with hostility simply for trying to talk with Native people while looking too White to be trustworthy. Prior positive engagement with specific individuals seems to carry little weight and be rapidly forgotten. The slightest misstep and, welp, "she's an evil White bitch, here to fuck over the Natives -- like they always are!" I'm not blaming people of color for feeling that way. I'm just saying that's the reality you are up against. As someone who spent some years homeless and got a fair amount of "help" offered of the "God, you clearly are an idiot causing your own problems and need a swift kick in the teeth as part of my so-called help" variety, I really sympathize with such reactions. White people often have little to no understanding of the lives of people of color and little to no desire to try to really understand because really understanding it involves understanding systemic racism in a way that tends to make Whites very uncomfortable. It veers uncomfortably close to self-accusation to honestly try to see how the world is experienced by such people.
- dontreact 6y agoNote that lung cancer screening is covered my Medicare and thus already free for anyone over 65 who smoked a pack a day for 30 years (or equivalent aka more in less time). My understanding is that there are many reasons that screening is not deployed more widely but the fact that it requires a 40 minute discussion with a physician, and those physicians in communities in need have very limited time. Then there is the issue of getting people to show up and take part in preventative care which is itself tricky. In any case, it was not something we were in a position to do much about as a small AI research team. Where I work now there is also a focus on trying to address this issue by reaching out to more hospitals to gather more diverse data, but there are still a lot of roadblocks to sharing data we have to work through and it’s a very slow process.
- 908B64B197 6y ago> The only data we were able to get at the time was mostly white patients. We talked to many hospitals but many were/are reluctant to share anonymized data for research I joked to someone from a small country with public healthcare that the best thing they could do was release as much anonymized high quality data and essentially get every machine learning algorithm tuned for their population for free. It's like adding your code to a popular CPU benchmark.