10 ms·
Towards accurate differential diagnosis with large language models
- guender 3y agoYikes, do I read this correctly and the LLM alone outperforms clinician + LLM?
- lucubratory 3y agoYeah, it's a bit awkward.
- ugh123 3y agoIt also outperforms clinician + googling which the study observed (many doctors use google search to help with diagnosis, terminology, recent studies, etc)
- chrstfer 3y agoHell, the LLMs themselves outperformed clinician plus LLM...
- fnordpiglet 3y agoThis has come up before in similar contexts. Model based decision decisions tend to do better than clinicians alone or clinicians with the model based decision system in general. Despite this it’s been effectively impossible to deploy such systems in a clinical setting due to many issues but one is that clinicians aren’t willing to cede their ground and patients aren’t willing to believe the machine over a human.
- caycep 3y agogranted, I feel like the training of some physicians also doesn't make them that good at diagnosis. Over time, intuitively understanding the narrative history of the illness/symptoms is key; that being said, nowadays, this sort of thing seems scattershot in some younger physicians. Med schools tend to try and rejigger their curriculum frequently to try and justifiably make the experience friendlier and w/ less workload for the students, but sometimes, it backfires in terms of this. I had an older school professor who liked to remark, in his day, no such thing as a differential diagnosis, just the right one and a bunch of wrong ones.
- 0xDEAFBEAD 3y agoIs there any way for me to get ahold of one of these models for my own private use?
- rafaelero 3y agoI don't see how those models couldn't be used to at least offer a second opinion. They are pretty cheap to run. I doubt patients would complain about that.
- rafaelero 3y agoThat surprised me as well. The case for humans being merely assisted by AI may not be the strongest one. It could be that they work better without human input, which is at the same time amazing and a bit frightening.
- Terr_ 3y agoJust to make analysis more complicated, not all errors in are created equal: It may be better to have lots of safe errors versus a few huge ones.
- kelseyfrog 3y agoIs that across all subspecialties? I hear internists talking about +LL and -LL all the time and I'd hope that reflected some higher degree of rationality when reasoning out diagnoses, but perhaps not?
- twobitshifter 3y agoMakes sense. In many applications the ‘human in the loop’ will become the weakest link. Imagine giving alpha go to players as a suggested move assistant. Even the best players would have seen alpha go’s suggestions and thought they were an error. Can we eventually trust AI even if we can’t understand its reasoning?
- chrstfer 3y agoThey'll be able to explain their reasoning......
- twobitshifter 3y agoExplaining and understanding the explanation are different.
- westurner 3y agoDifferential diagnosis > Machine differential diagnosis: https://en.wikipedia.org/wiki/Differential_diagnosis https://en.wikipedia.org/wiki/Differential_diagnosis CDSS: Clinical Decision Support System: https://en.wikipedia.org/wiki/Clinical_decision_support_system https://en.wikipedia.org/wiki/Clinical_decision_support_syst... Treatment decision support: https://en.wikipedia.org/wiki/Treatment_decision_support https://en.wikipedia.org/wiki/Treatment_decision_support : > Treatment decision support consists of the tools and processes used to enhance medical patients’ healthcare decision-making. The term differs from clinical decision support, in that clinical decision support tools are aimed at medical professionals, while treatment decision support tools empower the people who will receive the treatments AI in healthcare: https://en.wikipedia.org/wiki/Artificial_intelligence_in_healthcare https://en.wikipedia.org/wiki/Artificial_intelligence_in_hea...
- deleted 3y ago[deleted]
- rrsp 3y agoI wonder if they made any effort to check whether the NEJM case studies that this whole study is based are in the PALM-2 training dataset.
- resters 3y agoIf it's not already obvious, LLMs are going to be doing most of the mental work currently performed by doctors, lawyers, accountants, etc. I have already nearly stopped using Google search for anything, in favor of GPT-4. GPT-4 has helped me very quickly prototype things that I normally would have had to spend hours researching. GPT-4 has also created custom curriculum for me to help me learn various things for which I have struggled to find good books/tutorials online. There will always be many areas in which a solid human intellect and well-honed human judgment is still useful, but much of the less critical work will yield to LLMs. If we call the difference between GPT-3 and GPT4 1x, then I would expect to see a 2x-5x improvement in LLM capability within the next few years just based on how much great work has recently gone into shrinking really big models so they can run on smaller hardware. LLMs are not digital human minds, they are simply very good information synthesizers. Information synthesis happens to be what most white collar professionals get paid to do with their brains.
- hyperliner 3y ago[dead]
- doctorpangloss 3y ago> If it's not already obvious, LLMs are going to be doing most of the mental work currently performed by doctors, lawyers, accountants, etc. While I don't know you personally and I'm not talking about you specifically, "just" feeling very excited about something doesn't Make This About You, the prognostications don't actually help you Be a Part of This. This is the same energy as being really into COVID-19 (what was it called, "corona-scrolling?"), the gambling-fueled crypto boom, the retail stock trading booms. It's kind of adjacent to the fallacy Feelings that are More Strongly Held Are More Valid. It's like you could use these things as a barometer for the health of a social media channel. Anyway, I'm sure there's a name for this part of the hype cycle - tapping into how people want to be a part of the number one biggest, hottest trend by, at the very least, talking about it all the time, trying to compete against each other by sucking the most air out of all rooms through ever-greater hyperbole. One thing's for sure: there's some guy who found some massive low hanging fruit in some neural network designs by carefully choosing brackets in chained matrix multiplications. That's one guy, and he definitely Made It About Him in a way that makes sense. How many people on Earth have enough knowledge to see things like that, maybe 1,000-10,000? It's so exclusive, in a sense. I guess my point is, you're pretty far off the mark. In the interest of curiosity: Standard practice in medicine is to examine the patient before giving an opinion, a big roadblock for even the most competent of multi-modal models. Also, most people are biased in that they are intimately familiar with their own lived health problems, so they fill in tons of blanks, discounting the value of inquiry during a consult, whereas doctors may see you for less than a few hours a year, LLMs much less so.
- deleted 3y ago[deleted]
- rickysahu 3y agoWe are doing similar work at GenHealth.ai and getting sota results on some evals (not yet published). Our approach is very different from LLMs in that we are using a medical coding vocabulary and we are training transformers on actual patient histories. We have an API if anyone here wants to build on it. Oh and we are hiring
- doix 3y agoThis doesn't really surprise me given my experience with doctors. If it's a common problem then doctors do fine. As soon as you're not the common case or have something somewhat rare, it becomes a bit of a crap shoot, at least in the UK when you deal primarily with your GP (general practitioner). I have a friend with Crohn's who was feeling low energy. I was a gym bro at the time and convinced him to take a testosterone test (because all problems re caused by low T when you're a gym bro). His doctor wouldn't even entertain the idea, saying he's a young man and it's very unlikely that he'd have low T. He did the test privately and his T is significantly below normal. If you Google, there are actually many papers showing correlation between Crohn's and low T. I bet an AI would find it. Similarly, doctors missed my mum's recent cancer diagnosis. She also had factors that would make her more susceptible to breast cancer, googling finds many papers that show causation. The problem is that those things aren't extremely common and haven't made their way to NICE guidelines or whatever GPs use. Not that I'm blaming doctors, they have 10 minute appointments and don't have time to do anything. I'm sure AI would recommend significantly more lab tests which would put even more pressure on the NHS.
- corethree 3y ago[flagged]
- acover 3y agoYou can make the right call but still be wrong sometimes. There's a cost to tests of rare diseases.
- corethree 3y agoThen let the person who is wrong foot the cost. The cost isn't just the cost of the test. It's the cost of what occurred because the doctor decided to block a test.
- evgen 3y agoNo one is preventing you from going to a private clinic and paying for the test yourself. The doctor is not 'blocking' the test or preventing you from getting it, they are just not authorising it. Sounds like you have a problem with either your insurance provider or your wallet.
- hackerlight 3y agoIn poor countries there aren't enough doctors. Billions of people don't have access to healthcare and that won't change until the country becomes wealthier. Public healthcare that covers everyone can't happen because the reality of poverty makes it impossible. For these billions of people, the valid comparison isn't human doctor vs AI, it's no treatment vs AI. I am extremely optimistic about the positive role of AI in medicine for this reason.
- vidarh 3y agoThe irony is that we may see another leapfrog moment similar to analog cellphone networks and Internet infrastructure (where many poor countries largely skipped analog and modems because they had their boom later), where it will be easier to deploy AI for these purposes in countries where the existing healthcare system is underfunded and lacking in political clout than in richer countries even for situations where it delivers better outcomes.
- gcanyon 3y agoSure, but how will they train in an appropriate level of snark, abuse, and sexism to convey authority? (I miss House and hate how it ended)