17 ms·
AI models that predict disease are not as accurate as reports might suggest
- hdhdhsjsbdh 4y agoWhile the notion of treating these systems as “sociotechical” rather than purely technical is probably a good move wrt actually improving people’s lives, I can say from my own experience in academia that there are still way too many academics working in this field who don’t think it’s their problem. I’ve personally raised these types of issues before and been told “we’re computer scientists, not social scientists”, as if “social scientist” is a derogatory term. The biggest impediment here is, in my opinion, overcoming the bloated egos of the people who think the social impacts of their work are somehow out of scope. All is well as long as you can continue to publish.
- JHonaker 4y agoPreach. There are way too many people that conflate MSE or other abstract technical measurements of model performance like they actually represent the impact any model has on a problem. Even if we could somehow perfectly predict an actual realization instead of a conditional expectation that still forgets to ask the question of why we predicted that. Are we exploiting systemic biases, like historically racist policies? Almost definitely (unless we've consciously tried to adjust for them, and still we've probably done that incorrectly). I've become much less interested in models that basically just interpolate (very well I might add), and more in frameworks that attempt to answer why we see particular patterns.
- Grothendank 4y agoColor me, personally, surprised. Between publication bias and the general public ignorance of AI and its evolving capabilities, and over a decade of results in AI health being overblown before transformers, how could we have predicted that post-transformer results in AI health would continue to be overblown?
- photochemsyn 4y agoNot that surprising. AI learning seems to do best with fairly predictable systems, and when it comes to individual outcomes in medicine, there's a lot of mystery involved. A group of people with similar genetic makeup and exposure history to carcinogens or pathogens won't all respond identically - some get persistent cancers, some get nasty infections, and some don't. For example, training an AI on historical tidal data would likely lead to very good future tide timing and height predictions, without any explicit mechanistic model needed. Tides have high predictability and relatively low variability (things like unusual wind patterns accounting for most of that). In contrast, there are some current efforts to forecast earthquakes by training an AI on historical seismograph data, but whether or not these will be of much use is similarly questionable. https://sciencetrends.com/ai-algorithms-being-used-to-improve-earthquake-forecasting/ https://sciencetrends.com/ai-algorithms-being-used-to-improv...
- deltree7 4y agoYet. One thing media consistently gets wrong is the rate of innovation that is happening. Media also doesn't have access to state-of-the-art models, only from the trigger-happy startups too eager to release half-baked version. It's akin to downloading Image Generation tools from the App Store and concluding that's state of the art
- bpodgursky 4y agoIt baffles me that people can watch the trendline of "Job X can be automated in 40 years" (5 years ago) "Job X can be automated in 10 years" (2 years ago) "Job X can be automated in 5 years" (1 week ago) And feel comfortable poking holes in the AI models, pointing out where it fails. Obviously? But nobody 3 years ago thought that graphic design or creative writing was on death's row either. You have to spend a modicum of effort looking at how predictions have evolved over the past couple years, but once you do, it's very clear that mocking current AI systems makes you look like a clown.
- ckemere 4y agoThere's also the timeline that: "Radiology will be automatized in 5 years" (10 years ago) "Radiology will be automatized in 5 years" (5 years ago) "Radiology will be automatized in 5 years" (last year) or "Full self driving will arrive within 5 years" (5 years ago) "Full self driving is still a ways off" (last year) Assuming you're referring to generative models, I don't think that anyone (knowledgable) thinks that graphic design or creative writing are on death's door. They might change with new tools, but skilled practitioners are still required. That's basically the point of the article.
- chatterhead 4y agoWe are 18 years from the DARPA Grand Challenge and none of the vehicles finished. Do you think a self-driving car can make it from LA to NYC by itself now? What do you think 2040 AI will look like?
- lostlogin 4y ago
- dm319 4y agoAs someone who works in healthcare, so much of what I read about AI makes me think that the people who are enthusiastic about healthcare AI don't have much experience doing it. The scenarios rarely seem to fit with what I'm actually practicing. Most of medicine is boring, it is largely routine, and if we don't know what's going on, it's because we're not the right person to be managing the patient. Most of my time is spent talking to people - patients, colleagues, family. I explain the diagnosis, I talk about the plan, I am getting ideas of what the patient wants and values, and then actioning it. I spend very little of my time like Dr House pondering what the next most important test to perform is for a patient who is confounding us.
- ericmcer 4y agoThat scenario sounds like it lends itself more to AI automation than a Dr. House type one.
- kbenson 4y agoI don't know, compassion and understanding and nuanced understanding of individual desires when talking to someone is not what I associate AI with in my mind, but being able to assess sociological and cultural taboos and try to what a patient actually wants rather then what they might initially express seems like something I good doctor would get to through explorative conversation.
- junipertea 4y agoMaybe removing a human from the equation would lead to more honest outcome? E.g. people google all sorts of issues more earnestly than they would describe it to the doctors. The bottleneck would be properly understanding what the user intends, which might be out of reach.
- ben_w 4y agoIndeed. Language has been historically difficult for AI, but I think it's even tougher here — language is less and less reliable the further we get from a shared experience, and this is a problem when describing our experiences of our own bodies, and much worse when describing our own minds. For example, when I was coming off an SSRI, I was forewarned that I might get a sensation of "electric shocks"; the actual experience wasn't like that, though I could tell why they chose to describe it like that. How different is the tightness in the chest during a heart attack from the tightness in the chest from exercising chest muscles? I have no idea how doctors, GPs, and nurses manage this, though they seem to have relatively little trouble.
- johannes_ne 4y agoI recently published a paper, where we explain how an FDA approved prediction model, build into a widely used cardiac monitor was developed with an incredibly biased method. https://doi.org/10.1097/ALN.0000000000004320 https://doi.org/10.1097/ALN.0000000000004320 Basically, the training and validation data was engineered so an important range for one of the predictor variables was only present in one of the outcomes, making perfect prediction possible for these cases. I summarize the paper in this Twitter thread: https://twitter.com/JohsEnevoldsen/status/1561641153899929601?t=jVcYm9J9wGT1F-Y_LpC9lQ&s=19 https://twitter.com/JohsEnevoldsen/status/156164115389992960...
- baxtr 4y agoSorry for asking, but how is this relevant to the article?
- NovemberWhiskey 4y agoSorry for asking, but how is it not?
- baxtr 4y agoDo you agree that it’s ok to pose a question whenever you don’t understand?
- csallen 4y agoIronically, that's exactly what NovemberWhiskey is doing here :)
- ShamelessC 4y agoI’m not sure where you got this form of communication where you respond to everything with a question, and I assume you mean well, but it comes across as patronizing and de-humanizing to try to follow these “rules to winning arguments passively”, or whatever it is. Indeed, the confusion here is (I think) because your first comment > Sorry for asking, but how is this relevant to the article? Sounds accusatory. Please don’t respond to this with a question.
- drtgh 4y agoMy humble opinion; AI is supposed to be the acronym for artificial intelligence, but marketing has usurped it to refer to machine learning, which is nothing more than a neo-language for defining statistical equations in a semi-automated way. An attempt to dispense with mathematicians to develop models. What amount of energy is necessary for an event to be reflected in a statistic? You have a box of 2x2 meters with balls of data, and a string with a diameter of 1 meter with which to surround the highest concentration of balls possible, and those that remain outside, there they stay. Statistics and lack of precision are concepts that go hand in hand (someones say even it is not an science).
- jfghi 4y agoHaving built models, I’d claim that it’s art based upon science, perhaps not too different than engineering a building. At every stage there are decisions to be made with tradeoffs. Over time, the resulting model could be invalidated or perhaps perform better. It’s remarkably difficult to approach or even define a “best” model. What’s most peculiar to me is that somehow AI is becoming more distinct from math or stats and that there’s a notion by running pytorch one is able to play god and create sentience.
- spywaregorilla 4y ago> My humble opinion; AI is supposed to be the acronym for artificial intelligence, but marketing has usurped it to refer to machine learning, which is nothing more than a neo-language for defining statistical equations in a semi-automated way. Sure. Hardly controversial. > An attempt to dispense with mathematicians to develop models. What...? No. Definitely not. > What amount of energy is necessary for an event to be reflected in a statistic? You have a box of 2x2 meters with balls of data, and a string with a diameter of 1 meter with which to surround the highest concentration of balls possible, and those that remain outside, there they stay. Statistics and lack of precision are concepts that go hand in hand (someones say even it is not an science). I have no idea what this is saying. It sounds like you're shitting on statistics all of a sudden, which is weird, given that you seemed to favor mathematicians in the first part.
- 4y ago
- tensor 4y agoThis is entirely unsurprising and has a very simple solution: keep adding more data. Our measurements of the accuracy of AI systems are only as good as the test data, and if the test data is too small, then the reported accuracies won't reflect the true accuracies of the model applied to wild data. Basically, we need an accurate measure of whether the test data set is statistically representative of wild data. In healthcare, this means that the individuals that make of the test dataset must be statistically representative of the actual population (and also have enough samples). An easy solution here is that any research that doesn't pass a population statistics test must be up-front declared to be "not representative of real word usage" or something.
- blackbear_ 4y agoFrom the article: > Here’s why: As researchers feed data into AI models, the models are expected to become more accurate, or at least not get worse. However, our work and the work of others has identified the opposite, where the reported accuracy in published models decreases with increasing data set size.
- spywaregorilla 4y agoThat's not a contradiction per se. It's easier to get spurriously high test scores with smaller datasets. It does not clearly demonstrate that the models are actually getting worse.
- dirheist 4y agoBut if diagnosis are multimodal and rely upon large, multidimensional analysis of symptoms/bloodwork/past medical history, wouldn't adding more dimensions just increase dimensional sparsity and decrease the useful amount of conclusions you are able to draw from your variables? It's been a long time since I remember learning about the curse of dimensionality but if you increase the amount of datapoints you collect by half you would have to quadruple the amount of samples you have to retrieve any meaningful benefit, no?
- planetsprite 4y agoThe solution to failures of AI in heathcare is transparency of data. OpenAI's models work because they have virtually unlimited data to train on. The scale of training data for doctor bots is one millionth the size. Different countries, organizations, universities need to be as open as possible sharing and collaborating, realizing improvements in medicine benefits all of humanity with almost no downsides.
- dirheist 4y agoThere should be a standardization committee tasked with standardizing the collection of anonymized, semi-synthetic medical data from hospitals/hospital networks. It seems like so much research is just locked up in the IMS systems the hospitals use for their patients and that never see the light of day.
- dekhn 4y agoYou cannot imagine just how deep the medical data rabbit hole goes. Already plenty of institutions have semi-standardized their collect and do multi-hospital (typically research hospitals) aggregation. Whether this data is any good as training data for supervised or unsupervised algorithms is really questionable.
- potatoman22 4y agoI read about this researcher using GANs to create synthetic patient data, cool stuff. She had a problem she couldn't solve though: how do you validate that your synthetic patients look real?
- tomrod 4y agoTechnical (Honest) Solution: two holdouts 1. Involved in the build process 2. Never touched until paper metrics are being written, only run once Realistically, unlikely to occur however due to the incentives causing publication bias.
- jerpint 4y agoThird (better) option: have a regulating body have a separate, undisclosed test set. If you can't beat it, you can't deploy your model. If you can beat it, you still need to have your models peer reviewed and scrutinized
- tomrod 4y agoThis sounds simple yet I expect data governance will be the bottleneck.
- cm2187 4y agoso the models that fail this one test never get published and the models that succeed get published. And all you have done is to publish a model that predicts that particular history, in other words data fitting.
- tomrod 4y ago> the models that fail this one test never get published Not necessarily -- for comparison purposes one should include, and a negative result is not a bad outcome. But consider Edison's light bulb. Most of the failures don't matter, and a few might have interesting properties to reconsider or tweak down the line. But the major one folks care about is the one that worked. > data fitting Yes, models trained on data and not logic alone fit data.
- dr_dshiv 4y agoI’m shocked. SHOCKED.
- chatterhead 4y ago'Brunelleschi had just the solution. To get around the issue, the contest contender proposed building two domes instead of one — one nested inside the other. "The inner dome was built with four horizontal stone and chain hoops which reinforced the octagonal dome and resisted the outward spreading force that is common to domes, eliminating the need for buttresses," Wildman says. "A fifth chain made of wood was utilized as well. This technique had never been utilized in dome construction before and to this day is still regarded as a remarkable engineering achievement.' Brunelleschi was not an engineer he was a goldsmith. AI will advance in the same way architecture did during the Renaissance. By those with the winning ideas not with the right credentials. https://science.howstuffworks.com/engineering/architecture/brunelleschis-dome.htm https://science.howstuffworks.com/engineering/architecture/b...
- srvmshr 4y agoI worked in healthcare ML solutions, as part of my PhD & also as consultant to a telemedicine company. My experience in dealing with data (we had sufficient, and somewhat well labeled) & methods made me realize that a lot of the prediction human doctors make are multimodal - and that is something deep learning will struggle for the time being. For example, say in detection of a disease X, physicians factor in blood work, family history, imaging, racial genealogy, general symptoms (like hoarseness, gait, sweating etc), even texture & palpitations of affected regions sometimes before narrowing down on a set of assessments & making diagnostic decisions. If we just add in more dimensions of data to model, it just makes the search space sparser, not easier. Throwing in more data will likely just fit more common patterns & classes well, whereas a large number of symptoms may be treated as outliers and mispredicted. We humans are incredibly good at elimination of factors & differential diagnosis. The findings don't surprise me. There is much more work needing to be covered. For straightforward, and conditions with limited, clear cut symptoms they are showing promising advancements, but it cannot be trusted to wide arrays of diagnosis - especially when models don't know what 'they do not know'.
- dekhn 4y agoare you really sure the doctors are doing a better job when they go through the motions of incorporating a wide range of data? Or do we just convince ourselves they're better? I suspect we massively underestimate the amount of misdiagnosis due to incorrect analysis of data using fairly naive medical mental models of disease.
- srvmshr 4y ago> Are you really sure the doctors are doing a better job when they go through the motions of incorporating a wide range of data? Or do we just convince ourselves they're better? Personal story: I was diagnosed with a rare genetic disease in 2019. If I ran the symptoms through a ML gauntlet, I would be sure they would cancel each other out or make little sense. Chest CT (clean), fever (high), TB test (negative), latent TB marker (positive), vision difficulty (Nothing unusual yet), edema in eye socket (yes), WBC count (normal), tumors (none), hormones (normal) & retina images (severely abnormal) My condition was zeroed in within 5 minutes of a visitation to a top retina specialist, after regular opthalmologists were in a fix about two conflicting conditions. This was differential diagnosis based even though genetic assay hadn't returned yet, which also later came in favor. I cannot overemphasize enough how good human brain is in recalling information & connecting the sparse dots to logical conclusions (I am one of 0.003% unlucky ones among all opthalmological cases & the only active patient with that affliction in one of the busiest hospitals in the country. My data is part of the 36 people in a NIH study & opthalmo residents are routinely called in to see me as case study when I go for follow up quarterly).
- rscho 4y agoSurprise, surprise. People hugely overestimate the data retrieval capabilities of healthcare systems. And if you really put clinical 'AI' systems to the test in day-to-day settings (which is in fact never done), results would be much, much worse. Shit data in, shit prediction out.
- potatoman22 4y agoLook up prospective validation or clinical impact trials. These are validated in day-to-day settings somewhat often, though not as often as they should be.
- rscho 4y agoI'm a professional clinical researcher. Those 'validations' would be comical if the results, when witnessed in person, weren't so sad.
- potatoman22 4y agoThey are validated though. Not as much as they should be, but it happens. IMO the real blocker is hospitals don't want to/cant spend resources on these trials.
- bitL 4y agoOK, AI is bad but compare it to human doctors/radiologists that are often worse. I still remember stats from some X-ray detection where AI diagnosed with 40% accuracy and the best human doctors with 38% accuracy (and median human doctors with 32% accuracy). Now what are we supposed to do?
- pcurve 4y agoCan you cite the source? Is it not possible to improve the 40% rate by AI? Obviously someone eventually figured out the 100%
- dmurray 4y agoThey might have "figured it out" by cutting the patient open.
- yrgulation 4y agoOh god (science for some of us) the same kind of logic for defending tesla’s fsd system. Both crappy and dangerous, but with cult like following.
- naijaboiler 4y agoAlmost any paper like this is worthless. It makes for good headlines. But it does nothing to further medical care
- fasthands9 4y agoIt is still unclear to me exactly what data they were looking at/referring to in this article. If you take into account bloodwork, family history, demographics, etc. then it seems like you are still only getting a few dozen data points. At this scale it seems like traditional statistics or human checks for abnormalities are going to be about as good. Although I personally know very little (apologies for conjecturing) it does seem like there could be a lot of uses for AI for specific diagnosis. For example, when they take your blood pressure/heartbeat they only get data for one particular moment where you are sitting in a controlled environment. I would think if you had a year's worth of data (along with activity data from an apple watch) you might be able to diagnose/predict things that traditional doctors/human analysis could not. I would also imagine anything that deals with image analyzing (like looking for tumors in scans) will be vastly better with computer AI systems than humans.
- naijaboiler 4y agoLook simple rule based algos can be just as effective. The hard part is not diagnosis. It's getting the human to comply with treatment
- bjt2n3904 4y agoThis is what freaks me out about AI. People will use it for years in various fields, and one by one, after a decade or so of use, they'll come to find it was complete garbage information, and they were just putting their trust in a magic 8 ball. But the damage is already done.
- amelius 4y agoSame with self-driving cars. State of the art AI-based classification has an accuracy of 90%. Even if we can get it to 99%, that's still 1% error. Now imagine a car taking hundreds of decisions in a single ride.
- bmh100 4y agoThe issue with data leakage can be handled through k-fold cross-validation, in which all of the data takes turns as either training data or test data.
- cm2187 4y agoSo a model calibrated on a backtest says nothing about its predictive capacity. Who would have thought? Well, at least I think anyone who worked even a little bit in quantitative finance. The only way to validate a model is to make predictions and test if those predictions actually happen in a repeatable way, which in certain circles is refered to as "experiment". That's why I distrust any model built purely on backtested data unless they can be shown to predict something else than history. And AI is not the only area that blindly trusts those kind of models.
- yrgulation 4y agoSo many in ai chasing software solutions when the problem is hardware. Limited power means limited learning. Mix lab grown neurons with software and you have a wining proposition.
- rvz 4y agoOf course. No surprise there. Especially the ones made with 'Deep Learning'. At this point, Each time AI and 'Deep Learning' is applied and then scrutinised, it almost always concludes and tends towards pure hype generated by investors and output garbage unexplainable results from broken models. The exact same goes for the self-driving scam. 'AI' is slowing starting to be getting outed as an exit scam.
- trentnix 4y agoYet.
- rafaelero 4y agoAs always, AI performs poorly in the real world until they don't.
- dontreact 4y agoSweeping generalizations about AI models as a class of objects seem deeply uninformative to me. Perhaps I would even go so far as to say misleading. These days any undergrad can go on GitHub and have a model that does some diagnosis and I don’t understand why we have to group that together with multi-year efforts from groups of PhD’s and MDs working together to produce products. There’s obviously way more of the first than the second and so if you analyze this group as a whole it’s easy to draw the wrong conclusion about what AI as a technique is capable of. I can’t give too many details on specific examples I’m working on at Google but an example I’m not working on would be Caption Health who have an amazing AI-based ultrasound guidance product that has great prospective evidence, and several big fans in the clinical community. There are also several success stories using AI on pathology in order to target clinical trials. Can you imagine if someone made sweeping statements about webpages as if they were a coherent group of objects that you could sample from and deduce properties? “The information on websites is typically not as accurate as those websites claim”
- kimi 4y agoI feel it would be safe to say that "AI models that x are not as accurate as reports might suggest" for the current hype values of x.
- hef19898 4y agoAI seems to be a hyped solution in search for a problem. ML, sure, for some use cases. Regardless of where I encountered ML / AI / "expert systems in my life so far they have been sub-par to actual biological intelligence. In all cases they were really hyped.
- rickysahu 4y agowe are working on some transformer based models trained on millions of patients across a broad set of features. Well launch our first model into a private alpha in a few weeks. Basically it uses a seq2seq transformer XL and fhir data. https://labs.1up.health/ https://labs.1up.health/ Were hiring. U can email me directly at ricky@ We don't do any of the subset optimization or training for specific diseases or conditions. It's all or nothing like the way most llms work
- mpaepper 4y agoI work in machine learning for digital Pathology and I think the big problem here is the divergence between publishing papers and real life helpful models. What we see is that in the literature you often see a model trained on data from a single lab and gets crazy good results. However, apply it to a different lab (not in the paper of course) and it sucks. So what we do in practice for our models is to train on many different labs at the same time and have a hold out test set with labs not covered during training. That way you get a robust model which works well in practice (but doesn't have the 99.9% metrics reported in papers). The second thing is looking at what task to let the machine do. We typically go for tasks which are boring / repetitive for the doctor, but visualize the result, so the doctor double checks it before making the diagnosis. Still saves a lot of time.
- neuroguy123 4y agoI also work in this field and what we typically do is a lot of nested cross-validations to get some bounds on a model building process and some idea of how it would perform on repeated unseen data. Data leakage is always on our mind and we do our best at all stages to avoid that. We also train on data from many sites. It can be done and it can be done properly. As you say, it is always best to collect some completely naive test set to back up the model-building process. If you design your pipeline properly, the test set should fall within the bounds you got during cross-validation. It all depends on how much data you have and I think as long as you design your pipelines with that in mind and acknowledge limitations with smaller datasets, then the research is valid and useful.
- Dr3xus 4y agoDo you have any specific example use cases where you're seeing success? I'm asking this both with regards to ML/prediction related work and the "boring" work. I'm working on a direct to consumer digital pathology service with a clinical background as opposed to data science/ML. Really curious as to what type of new services we could try invest in to improve what we offer.
- SimonV1235 4y ago