23 ms·
I’m a doctor: Here’s what I found when I asked ChatGPT to diagnose my patients
- maherbeg 4y agoAh yes, the "everyone lies" House M.D. problem
- SketchySeaBeast 4y agoNot even, ChatGPT, being an engine that figures out what's right by finding out what is average, is bad at understanding the atypical.
- operatingthetan 4y ago>is bad at understanding the atypical. To be fair I've found most doctors require a lot of convincing if your problem atypical as well.
- SoftTalker 4y agoAs they should. If you hear hooves, think horses not unicorns.
- operatingthetan 4y agoIf you google around there's tons of stories of people with legitimate novel illnesses that are ignored or misdiagnosed for years.
- SketchySeaBeast 3y agoWhich sucks, but can you imagine the additional burden our health care system would be under if they had to test for every possible thing a condition could be every time you went in? I wonder what all the edge cases of something presenting like a headache could be.
- salawat 3y agoZebras, deer, moose, and giraffe, bulls of many types, all exist. I'd say if you hear hoofbeats, contemplate what hooved animal would be appropriate to the situation and try to eliminate the easy ones. Lord knows, you might find a Zebra escaped from the zoo, or a bull from someone's field from time to time. In short, life is messy, and taking any aphorism to the death will lead to just that.
- ChatGTP 4y agoIt’s funny because it’s almost the exact same problem I have with using it professionally for writing software.
- throwbadubadu 4y agoSeconding.
- deleted 4y ago[deleted]
- sixothree 4y agoI had an opportunity today that made it useful. I was trying to find a reliable method of counting the decimal places in a double (64 bit float). I couldn't help but feeling like the responses were not quite informed and possibly dangerous. Chat GPT provided a solution, one that appeared better than most of what I had seen in the previous 15-30 minutes. I asked it twice to ensure safety and it improved its response. I then asked it to explain a particular choice and it was thorough enough for me to feel comfortable. In the end I feel like it understood its reasoning better than some of the options I saw on SO. This was GPT-4 and a fairly simple problem that was benefited by its understanding of the double type.
- yieldcrv 4y agoYeah it bullshits a lot even on single liners Its a faster and better stackoverflow for me, which is a big value add because the community and moderation aspect of SO is absurd I love when it tells me about libraries and resources that I didn't know existed, when I didn’t necessarily ask the likely followup questions yet Break big problems into smaller problems and let it tackle them
- zzzeek 4y ago[flagged]
- jjcon 4y ago> because this doctor is not practicing in Texas where such a procedure might get you arrested https://texas.public.law/statutes/tex._health_and_safety_code_section_245.002 https://texas.public.law/statutes/tex._health_and_safety_cod... Please don't spread misinformation, there is enough confusion out there already. Texas law specifically allows for the removal of ectopic pregnancies.
- zzzeek 4y agothese specific provisions are insufficient and women are having harrowing health care experiences in Texas nonetheless due to doctors delaying or denying care out of fear of prosecution: https://www.texastribune.org/2022/09/20/texas-abortion-ban-complicated-pregnancy/ https://www.texastribune.org/2022/09/20/texas-abortion-ban-c...
- jjcon 4y agoExactly, due to tons of misinformation about the law like the above - it is causing tons of problems - we don't need to further spread it.
- zzzeek 3y agoyou are making excuses for fascists who are intentionally terrorizing women and I find it offensive. Denying care is not the kind of thing you do because you saw a rumor. these laws are designed to be vague and terrorizing and they are having the intended effect of punishing women for having sex. you are spreading disinformation that this is not the purpose of these laws. headlines about care being denied in Texas due to these laws are myriad. My Hacker News comment is not the cause of this. https://www.chron.com/news/houston-texas/article/Texas-abortion-law-hospitals-clinic-medication-17307401.php https://www.chron.com/news/houston-texas/article/Texas-abort... https://www.businessinsider.com/texas-abortion%20hospitals-refusing-treat-serious-pregnancy-issues-fear-report-2022-7 https://www.businessinsider.com/texas-abortion%20hospitals-r... https://www.npr.org/sections/health-shots/2023/03/01/1158364163/3-abortion-bans-in-texas-leave-doctors-talking-in-code-to-pregnant-patients https://www.npr.org/sections/health-shots/2023/03/01/1158364... there are so many headlines. But no, it's my fault, not the fault of these fascist lawmakers who wrote these laws and know very well what they are doing. excuses.
- Turskarama 4y agoI'm not really sure what he expected here, ChatGPT was not trained to be a doctor, it is far more general than that. Asking ChatGPT for medical advice is like asking someone who is very well read but has no experience as a doctor, and in that context it's doing very well. He also brings up one of the most salient points without really visiting it enough: ChatGPT does not ask for clarification, because it is not a knowledge base trying to find an answer. All it does is figure out what character is statistically most likely to come next, it has no heuristic to know that there is a task it hasn't fully completed. This is the same reason ChatGPT cannot yet write programs by itself: in order to do so you'd need to specify the entire program up front (which is exactly what code is). As soon as we have agents that can do a proper feedback loop of querying a LLM consecutively until some heuristic is reached then the kind of AI doctors are looking for will emerge.
- izzydata 4y agoHe probably didn't have any expectations. Merely experimentation and observation and maybe a bit of word of warning after seeing the results for anyone thinking of using it in ways it shouldn't be.
- onlyrealcuzzo 4y ago> All it does is figure out what character is statistically most likely to come next To be pedantic, it's a token - not just a character - right?
- gpm 4y agoYes
- JacobDotVI 4y agoMy hunch is that this is exactly what he was expecting. There is a lot of hype around ChatGPT passing the medical exam and this exercise is a counter point to that.
- famouswaffles 4y ago
- frognumber 4y agoThere are two types of medical conditions 1) Those you see a doctor for 2) Those you don't The line depends on where you live. In a poor village, 100% might be the latter, while an executive in SFO will see a doctor for anything serious, but might not if they cut themselves with a kitchen knife. What's underrated is the ability to have basic medical care and information everywhere, all the time, for free. That can be casual injuries below the threshold of visiting a doctor (am I better heating or icing? immobilizing or stretching?), or those can be settings where there are no doctors. Even more, doctors (like AIs) make mistakes, and it's often helpful having a second opinion.
- jhgg 4y agoI am curious if GPT-4 would have performed better.
- famouswaffles 4y agoit definitely would. How much better is the question.
- preommr 4y agoIt's amazing that it was that effective... - It's a generalized language model; imagine how much more effective it would be with a specialized ai that used a variety of techniques that are better suited for logic and reasoning, while using llms to interact with patients. - It cost an order of magnitude less than the visit to a doctor. - The potential in being able to constantly monitor a patient - a point made in the post.
- intelVISA 4y ago> only reflecting back to me the things I thought were obvious — enthusiastically validating my bias like the world’s most dangerous yes-man. This is why it's exciting: we're seeing that awkward stage of impressive (for entry level/passing the bar) but still requires (expert?) supervision. Any worse and the novelty would wear off - any better and we'd be having (warranted) AI panic. Could go either way tbh
- rzzzt 4y agoI imagine development and training of these mythical non-LLM approaches are now placed on the backburner while the world is collectively enamored with eloquent virtual assistants inside microwaves and calculator apps.
- chimeracoder 4y ago> So after my regular clinical shifts in the emergency department the other week, I anonymized my History of Present Illness notes for 35 to 40 patients — basically, my detailed medical narrative of each person’s medical history, and the symptoms that brought them to the emergency department — and fed them into ChatGPT. It's quite shocking that the doctor would openly admit to violating HIPAA in such a brazen way. HIPAA is incredibly broad in its definition of protected health information: if it's possible to identify an individual from data even through statistical methods involving other data that a third party might already conceivably possess, it's considered protected. It's inconceivable that the doctor would be able to sufficiently anonymize the data in this capacity and still provide enough detail for individual diagnoses. There are processes for anonymizing data to disclose for research purposes, but they're pretty time-intensive, and no ED would allow a doctor to do it by himself, nor would they provide that turnaround in just "a couple of weeks". And the end results are a lot less detailed than what's needed for individual diagnoses like these. I really wonder what the hospital will say if and when they see this post. Given the timeframe and details described in the post, it's really hard to believe that they signed off on this, and hospitals don't take lightly to employees taking protected and confidential data outside their systems without proper approval. EDIT: It looks like this doctor works at a for-profit, standalone acute care clinic, rather than a traditional ED at a hospital, so my statement that hospitals don't take lightly to this stuff doesn't apply. The law still applies to for-profit standalone emergency care, but they tend to play fast and loose with these things much more than traditional health networks.
- tedunangst 4y agoAnd yet medical journals are filled with articles with sufficient detail that other doctors can even learn to make diagnoses from reading them.
- chimeracoder 4y ago> And yet medical journals are filled with articles with sufficient detail that other doctors can even learn to make diagnoses from reading them. This would be an apt analogy, if medical journals involved no oversight from the covered entity at which the patient presents, if there were no editorial intermediary, and if the entire publication timeline happened in weeks, allowing for no data redaction and review, rather than years.
- suddenclarity 4y agoIt does worry me what data people are sharing without seemingly much though. He claims it anonymised but I'm a bit sceptical when you input the medical history of 40 people. It's easy to slip up.
- jacquesm 4y agoWith a rare enough disease the anonymized file would still be enough to ID the patient given where the doctor is located.
- jeroenhd 4y ago"I fed my patients' medical information into this tool that promises to regurgitate it for others" is one headline I didn't expect to go down so easily. Running this stuff through an offline LLaMA instance? That seems fine, the software can't leak anything and doesn't retrain itself. But using ChatGPT? That simply cannot be legal. Stories like these make me distrust doctors. Very few of them seem to care about privacy outside of telling people I know about my medical issues. Nurses gossiping about patients is bad enough. I really don't want a future where I'm going to need to find a doctor that avoids recent technological developments because they're too uncaring or technically incompetent to not feed my most private information into some big tech company's algorithm.
- PaulKeeble 4y agoBut it also gets around the common misdiagnoses for chronic conditions. It has a great description of Long Covid and ME/CFS for example whereas your typical Primary care is going to dismiss that patient with a Psychology diagnosis as is happening daily across the entire western world. Its less biased but its not going to find the rare things especially where the patient has missed something important. Its a mixed bag just like it is with software. If you ask it to solve something simple it often does a decent job, but something complex and its confidently wrong. It doesn't show the self doubt of expertise that it needs to be a reliable tool yet it still requires the user has that expertise to be able to save time using it.
- gamesbrainiac 4y agoWhich version though? 3.5 or 4? It does not state this explicitly. There is a world of difference between 3.5 and 4.
- xiphias2 4y agoThis is a republication of an older article that was published just when ChatGPT 4 came out, and the date was changed. I personally had seen good and bad parts of diagnosing with ChatGPT 4, and what I would interested in is if the doctor tries using multiple questions and finds out how to use the tool well. I believe he could have improved the tool significantly if he puts in the time to experiment with it.
- ekidd 4y agoSome of the best performances I've seen out of ChatGPT are essentially "junior programmer" level. But it still requires clear instructions and close supervision. But GPT's training data includes GitHub, and it's used to power Copilot. It has arguably been trained to be a programmer. In less familiar domains, like law or medicine, GPT has presumably undergone very limited training and tuning. It's essentially an "internet lawyer" or an "internet doctor." In domains like this, it simply can't provide zero-shot professional results. Not with the current training data sets, and not with the current model performance. Of course, we have no idea how quickly this gap will be closed. It might be 6 months or it might be 6 years. The future is looking deeply weird, and I don't think anyone has even begun to think through all the implications and consequences.
- s0rce 4y agotraining on uptodate.com would probably be a good start
- deleted 4y ago[deleted]
- aflag 4y agoNot sure if I'd compare ChatGPT with a junior programmer. In my experience junior programmers tend to be builders. They will tend to code a lot of stuff and usually get reasonable results, but making some bad decisions that more experienced developers have already gone through. Inexperienced developers need supervision because otherwise they will just create heaps of code that will be hard to maintain later. ChatGPT just doesn't do anything on its own and will never follow through with anything. So, it doesn't really need supervision. I feel like it's more like a professor or a very senior developer. Someone you'll consult with when you're having trouble. Obviously, our best specialists are still better than the AI, but if the current technology is perfected, it'd expect it to replace the specialist and not the junior programmer. Which obviously is a bit of a bleak future from a software engineer career's perspective.
- ekidd 4y ago
- orcajerk 4y agoBack in 2004, took a seminar in college regarding Decision Support Systems and how they a manager or doctor could ask it a question and get a response to help them make a decision. Went to the doctor couple a years ago and he charged $300 to google search the symptoms. No thanks.
- intelVISA 4y agoYou're expecting a doctor to have all relevant medical knowledge permanently memorized? That's the equivalent of coding interviews on random obscure topics where you can't look anything up. Like a SWE their value is not perfect recall of every area of CS/medicine but ability to decipher arcane documentation into actionable outcomes.
- dekhn 4y agoIIRC my relatives who got medical degrees all commented on just how much memorization is involved.
- jeroenhd 4y agoThe $300 were not to Google the symptoms, they were to sift through the bullshit Googling symptoms will return that the doctor knows won't apply to you. Looking up symptoms without regard of likelihood is how you get "I have either the flu, stage 3 cancer, or drug-resistant super AIDS". Most tech support is little more than Googling the right question and going through the steps in the first or second result. Knowing what questions to Google and what answers won't apply is the reason you get paid for that stuff. I, for one, like my doctor to use tools to find possible diagnoses that she may have learned about 30 years ago but rarely ever come up, as long as the tools they use preserve my privacy.
- gilbetron 4y agoBeing a scifi geek and AI geek and neuroscience geek for pretty much the past 40 years, I've read countless predictions and scenarios and stories about society's response as "true" AI begins to emerge. So watching it play out for real is creating this bizarre sense of deja vu combined with fascination and frustration and also some anxiety. This article and the comments in this thread are right up that alley. I mean, can you imagine say 1 or 2 years ago saying we'd have a readily accessible system that you could feed it the symptoms a patient is experiencing (in any language!) and out would split a well-described explanation of the diagnosis (in any language!) around half the time? And now it is suddenly here, and so many people are like, "yeah, but it isn't perfect" or "we can't completely trust it" or "we have to fact check it because it can be wrong". It is all so weird and fascinating and just bigger than I can describe. But here we are!
- famouswaffles 4y agoIt's interesting for sure. There's also the fact that he didn't use the sota (GPT-4). I'm not saying that would be perfect. But i would be surprised if it wasn't big leap over 3.
- weaksauce 4y agoit wasn't mentioned anywhere in there which version of the model he was using... how are you so sure?
- famouswaffles 4y agohe references 3 when he talks about passing the exam. and i guess i doubt he'd evaluate 4 and not say anything about it.
- weaksauce 3y agothey didn't specifically mention 3 in the linked article or in the linked linked article and they ran both 3 and 4 against those exams so I'm not sure either way. probably was 3 but I can't be certain.
- sixothree 4y agoThe number of people in this country who don't even have a doctor is abysmal. I would think an AI doctor would be better than no doctor at all.
- mabbo 4y agoIt got 50% right with a general model. That same model can do a pretty good job at a thousand other skills. Imagine what it can do if it's trained for exactly this purpose.
- dekhn 4y agohttps://blog.google/technology/health/ai-llm-medpalm-research-thecheckup/ https://blog.google/technology/health/ai-llm-medpalm-researc... is an example of a more tuned model At some point people are going to notice that the best of these systems disagree with the experts who disagree with each other on the harder problems. That's one of the issues with AlphaFold- the accuracy of their predictions are so good, its predictions have metric values as good as the limits of the experimental technology used to create the golden labels.
- debacle 4y agoDevil's advocate, but maybe this doctor should be better at charting?
- bsuvc 4y agoThis is what I was thinking too. Often when I find a response to be lacking, I try again with a better prompt and get the kind of response I was looking for. Part of AI's effectiveness is in the human's ability to prompt well.
- m3kw9 4y agoWould be great if OPENAI can get hospitals around the country could anonymize some trainable format data
- deleted 4y ago[deleted]
- flir 4y agoThis is how I played doctor with ChatGPT. > Do you know what an Expert System is? [describes expert system] > I'd like you to emulate an expert system while I play the role of a patient. Expert systems ask questions in order to produce a diagnosis. I would like you to ask me a series of questions, and respond with a diagnosis when you feel confident. Alright, I'll do my best to emulate an expert system. Please keep in mind that I'm an AI language model and not a certified medical professional. For any serious concerns, it's essential to consult with a healthcare provider. Let's begin. What is your primary complaint or reason for seeking help today? I fed it a symptom my doctor had already diagnosed, and it did ok - it got it down to three possible causes, one of which was the correct one. All along the way it was warning me that I really should see a real health professional and it's just a chatbot. What really interested me is that I said "please emulate an expert system" and it did. Once upon a time, expert systems were an entire branch of AI, and here it is just emulating one off the cuff.
- TheRealPomax 4y agoAnd failing at it, if you had to help it at every step.
- sjducb 4y agoI suspected it would do better with good prompt engineering
- jcims 4y agoI know everyone scoffs at the concept of 'prompt engineer', but it really is an essential craft that we're going to have to come to terms with when interacting with large language models. Seeking suggestions on a more comprehensive prompt: https://sharegpt.com/c/sckAPvV https://sharegpt.com/c/sckAPvV Trying it out: https://sharegpt.com/c/LbpEIxi https://sharegpt.com/c/LbpEIxi
- flir 4y agoPlease see my "expert system" approach elsewhere in the thread. I felt it worked really well.
- dekhn 4y agoif you want more prompt engineers, just have kids and as part of their growing up, teach them to prompt engineer and make it mildly competitive. Some fraction of them will be better than nearly any current prompt engineer. children: the OG AGI
- Enginerrrd 4y agoI agree. I'm a civil engineer / project manager and so far I've been VERY impressed with chatGPT and, in particular GPT-4. However, a huge part of my job has always been translating vague desires into very precise specifications with constraints and expectations. Going further, it has often been my job to take those specs/constraints and then break them into chunks and feed them to junior staff who are often very smart, but lack domain specific context and knowledge. Giving them a bad prompt produces bad results. This article seems to be based largely on data collected with a rather poorly engineered prompt, IMO. He asked it a question that would be reasonable to ask a fellow physician. The problem is GPT is NOT a fellow physician with domain specific context and knowledge, and isn't aware of a bunch of implicit expectations they didn't realize they had. However, I actually think there's a really good chance that a better worded prompt would have scored a lot better here. This type of communication skill has always been hard for a lot of people, and will remain in high demand for a long time.
- AmericanChopper 4y agoI don’t scoff at it, but I do think it’s kinda funny. It’s essentially the same skill as being a “google search expert” in the sense that you need to be able to correctly understand the problem, and craft a good enough query to generate the answer you’re looking for. It’s always sort of been a tongue-in-cheek claim that googling skills are a valuable software engineering asset, and even though that’s a perfectly legitimate assertion, it’s entertaining to see it emerging at its own “engineering” discipline.
- scotty79 4y ago“Any chance you’re pregnant?” Sometimes a patient will reply with something like “I can’t be.” “But how do you know?” If the response to that follow-up does not refer to an IUD or a specific medical condition, it’s more likely the patient is actually saying they don’t want to be pregnant for any number of reasons. Funny how languages are ambiguous around "can't" and "don't want".
- seanp2k2 4y ago>If my patient in this case had done that, ChatGPT’s response could have killed her. Not if she lived in a state where there's no longer any legal treatment for ectopic pregnancy.
- ghiculescu 4y agoWhich states are you referring to?
- jppope 4y agoI think this doctor is forgetting about the other side of the coin... would chatgpt perform better than a really bad doctor?
- mattgreenrocks 4y agoI asked ChatGPT to write out a G major Ionian scale with three notes per string in guitar tablature notation last night. Mostly cause I was too lazy to do it myself. After 7 rounds of me fixing its mistakes, I gave up. It doesn’t really know what it is doing, so I can’t make forward progress. It put two notes on one string, repeated notes from a lower string on a higher, put the scale out of order, and forget previous corrections. Whatever hope I had of saving time was completely lost. I eventually realized the correct thing to do was either make my own charts or just practice them in F like they were made. I’m skeptical that scaling the model up will cause it to learn this, and I don’t consider this a very complex thing to learn. No, I didn’t try GPT4.
- stephendause 4y agoI have tried GPT-3.5 and 4. There is a marked difference. (I have used it do what I think of as simple but nontrivial programming tasks, asked it for recommendations for various products, etc.) GPT-4 still fails for me regularly but not nearly as often as GPT-3.5. I find it useful enough in my daily life to pay $20/month for. So if you did want to try it, you might be surprised.
- NiloCK 4y agoI don't know how much this sort of thing is frowned upon here, but I wrote an article about this scenario recently. http://paritybits.me/disposable http://paritybits.me/disposable tldr: the gpt services will eventually (maybe soon) recognize opportunities to write and run their own bespoke software to provide higher resolution outputs.
- nmfisher 4y agoI have one specific task where GPT-3.5 failed completely, but GPT-4 succeeded spectacularly (generating correctly formatted AutoRig Pro bone mappings in Blender from one armature to another). 4 still fails regularly on a lot on seemingly basic tasks, but it is a noticeable step up from 3.5. As they continue to scale it up, I suggest checking back in every few months to see if the newer versions perform any better.
- phoenixreader 4y ago
- pyrophane 4y agoTo me this just reinforces the notion that if you train a system like this to be a doctor it would be very effective.
- blastonico 4y agoI think that his point is: (1) don't use chatgpt for self diagnosis. Go see a doctor. (2) Doctors, chatgpt isn't ready or the right tool to help with your duties.
- rafaelero 4y agoAuthor didn't mention if he used GPT-3.5 or GPT-4.
- swader999 4y agoI'm thinking of parsing my wife's (vet) text books into a vector db format and then doing a search on those pdfs for relevant text to submit as part of a decent prompt. She'll tell me pretty quickly if this is useful or not.
- tejohnso 4y ago> about 8% of pregnancies discovered in the ER are of women who report that they’re not sexually active. This is the most surprising thing I read in the article.
- Atsuii 4y agoAs woman this doesn't surprise me at all. There is a lot of circumstances where a woman would not want to admit to an ER that they are sexually active; Religion, the sex was non consensual, they are with someone who doesn't know they are sexually active.
- lanstin 4y agoI think we will learn more about the biased sample that transcribed speech has vs private speech. I have just started thinking about this but there are huge area of speech that I use with my family and wife, and coworkers around the coffee room, that I would never consider putting into writing or a YouTube or whatever. And more types of things that I might say when with some younger folks while, trying to debug s production emergency but would not make it into the post mortem. The training I guess is not going to be able to benefit from these more private speeches, or maybe we will have to have people become convinced they are sentient and fall in love with them and share their whole humanity with the LLM.
- petilon 4y agoOn the other hand, I have multiple minor issues where doctors have not been able to offer a diagnosis (they just say "I don't know") and ChatGPT has been able to offer multiple possible diagnoses.
- egl2021 4y ago"I don't know" is exactly what I want my docs to say when they don't know.
- petilon 4y agoRight. And going forward, if they take help from ChatGPT they will have to say that less often.
- aix1 3y agoDo you really though? Wouldn't you want them to say "I don't know, but here are the next steps I think we should take to continue getting to the bottom of this (further tests; the doctor reviewing literature or consulting with colleagues; a referral to a different specialist etc)"?
- 627467 4y agoI wonder if you could prompt engineer your way to chatGPT to pretend to be a doctor and behave like one in, like, asking questions. Many already said, chatGPT is not optimized for any scenario. I don't doubt that training it for medical applications is already underway. I mean, flesh and bone doctors in many countries already behave as bots essentially reading/answering through a sequence of questions on a screen. I can definitely see most GP being replaced by bots of sometime or people who are actually trained to display empathy with patients.
- lamontcg 4y agoI'm finding more or less the same behavior with ChatGPT when it comes to programming problems. If I feed it some leetcoding-like problem it usually gets a pretty good answer. I used it to write some rust code to strip-underline-followed-by-trailing-digits off a string. The first guess it made though was to strip a substring, and it missed the fact that the slice it was using could panic. The third try at least passed my "peer review". It was useful because after a decade of using ruby my instinct is to reach for regexp captures, the solution it came up with is probably a lot faster and easier to read and avoids "now you have two problems". I tried to get it to help me eliminate an allocation caused by the capture of variables in a lambda expression in C# and it just started to aggressively gaslight me and break the code and claim it was fixed (very assertively).
- textninja 4y ago> my instinct is to reach for regexp captures, the solution it came up with is probably a lot faster and easier to read and avoids "now you have two problems”. I don’t write Rust but I think it’s best to trust your instinct here. “Now you have two problems” is a humorous quip and not practical coding advice. Using a regular expression to strip trailing digits from a string will surely result in code that is shorter and more readable than the alternatives, and it will probably be more correct too.
- lamontcg 3y agoYou might be entirely correct. I noticed a small perf issue with the backwards scanning and had it rewrite it not using rfind("_"). First time I prompted it, though, it didn't actually manage to fix the problem and gave me some mild gaslighting until I told it to explicitly remove the rfind() call. Now I think the result should be comparable to or beat a regexp, but it is getting quite a bit more complex: fn strip_suffix(s: &str) -> String { let mut idx = s.len(); for (i, c) in s.char_indices().rev() { if c == '_' { idx = i; break; } else if !c.is_ascii_digit() { return s.to_string(); } } if idx == s.len() { return s.to_string(); } if s[idx+1..].chars().all(|c| c.is_ascii_digit()) { return s[..idx].to_string(); } s.to_string() } .char_indexes().rev() there worries me a bit as well now that I look at it... haven't tested that at all.
- matt_heimer 4y agoI argued with someone (online) about a ChatGPT diagnosis recently. They have back pain and they've had an MRI but the scans (shared online) don't show any significate disk bulging and no herniations. They put their symptoms into ChatGPT and got a possible diagnosis of a disk herniation along with a treatment of a microdiscectomy. Despite the fact that a couple doctors have told them they don't need surgery they are convinced that they do. I understand that they are desperate for a solution to their pain but they are now doctor shopping until they can find someone willing to perform a procedure that ChatGPT suggested. People are already being misled by these systems.
- Waterluvian 4y agoYou said it: they’re desperate. People do considerably less rational things in a desperate attempt to make pain go away.
- recursive_loops 4y agoI've recently been using ChatGPT to learn about Linux networking stuff, including most recently how to setup a bridge connection for KVM that behaves the same as virtualbox's. Given how catastrophically wrong (confidently wrong at that) ChatGPT has been, I cannot even imagine the frustration that will be for doctors from people who don't understand how LLMs work and think they are "thinking."
- QuantumGood 4y agoWhen GPT starts to auto-incorporate the best yet known prompts, then we'll have a better idea of its potential. You must, must use the best prompts, of which many are not widely known, and some have not (of course) been discovered ... yet. Even with human experts, you must provide sufficient detail, and the expert must ask clarifying questions for differential diagnosis.
- jacobsenscott 4y agoThis is nonsense - a LLMs output is a pseudo random mishmash of its training data. How can there be a "best prompt" when the same prompt gives you different output each time, but there is only one correct output?
- syntheweave 4y agoCalibration of the response occurs through prompting. That is, if you ask GPT a question directly, it will give the most glib, unfiltered answer possible like an average Reddit user. But if you tell it to role-play a persona it will apply elements of that persona, causing it to filter itself. Or if you tell it "describe your reasoning step by step" it will proceed down a path with greater causality. Or if you tell it "here are the rules of a language, give a response to the prompt in this language" it will attempt to produce a response that looks like those rules, using the weighting associated with logical words like "must," "cannot," "if," etc. Measuring calibration is the problem now: we know some prompts do much better than others, but not how to optimize that in a general sense, to make an LLM always adopt the persona needed for the job.
- fwlr 4y agoPrompts constrain the possible output space. This is true in trivial ways: ask it to reply only in json. This is also true in slightly less trivial ways: ask for a “description of X”, a “short description of X”, and a “one-sentence description of X”. This continues to be true in increasingly more complex ways: Prompt it with a quote, a rating of that quote as 5/10 on complexity, and request it to give you two new quotes with ratings of 2/10 and 8/10 on complexity. Follow it up with a request for two more quotes at -5/10 and 20/10 ratings. Then try the whole process again but with metrics other than “complexity”: information content, eloquence, humour. In this way, if there is a single correct output, there is a corresponding single best input prompt that most tightly constrains the output space to the smallest space that still contains the correct output.
- drewcoo 4y agoWell there's an ethics lawsuit advertising to happen. I certainly don't want my docs handing my medical information to ChatGPT, even if they believe they've "anonymized" it.
- paraxion 4y agoI wonder if, instead of asking ChatGPT for a diagnosis, he could've got it to prompt for further questions he could ask? My thinking is that given the nature of LLMs of connecting related information, it might be a good way to figure out the gaps in the diagnostic process, rather than actually provide one.
- pcthrowaway 4y agoI was thinking the same thing, the author may be a great doctor, but not a great prompt engineer (perhaps even intentionally so, to justify their job) Instead of "Here are symptoms, what are possible diagnoses?" They could have tried "Here are symptoms, what are possible diagnoses, and what are some good questions an intelligent doctor might ask to be able to better diagnose their patient?"
- badcppdev 3y agoOr the doctor could feed their answer into ChatGPT before they give it to the patient and ask if there are any possible errors
- Xcelerate 4y ago“I’m employed as an [X]. Here is why machine learning won’t ever replace my job. Even once the point has been reached where all overwhelming evidence points to it being better than me in every regard, I’ll still summon an emotional appeal to that ‘magical’ human quality that machines will just never replace.” To be honest, I think I’d rather be friends with ChatGPT than most humans as it continues developing over the next decade.
- nextworddev 4y agoWhat’s interesting is how “AI hallucination” and “prompt hijacking” are supposed to be such horrific problems - when humans lie or make mistakes in recall all the time, not to mention get gaslighted into say awful things.
- sourcecodeplz 4y agoOmg stop with this ridiculousness ffs. I get and love AI but some areas should be off limits: doctors, judges, airplane pilots, train conductors... Soon enough no one will even know how to write, just read, because ChatGPT will write everything.
- aix1 3y ago> some areas should be off limits: doctors, judges, airplane pilots, train conductors I find this list very odd, especially given that we've had driverless train systems for a number of decades: https://en.wikipedia.org/wiki/List_of_driver-less_train_systems https://en.wikipedia.org/wiki/List_of_driver-less_train_syst...
- textninja 3y ago> Soon enough no one will even know how to write, just read, because ChatGPT will write everything. Nonsense, writing is easy! Just dictate some rough instructions to a GPT agent and copy/paste its response. I’m being facetious, of course - writing is thinking, so I don’t think it’s necessarily going anywhere, though AIs can obviously augment or replace a lot of the busywork. Where ChatGPT is used to generate content absent thoughtful prompting, the stuff it spits out will largely be regarded as spam.
- ftxbro 4y ago"It diagnosed another patient with torso pain as having a kidney stone — but missed that the patient actually had an aortic rupture. (And subsequently died on our operating table.)" Wow imagine if the AI had been used in an unquestioning way. Someone could have died!
- prirun 4y agoUsing AI to find patterns across many patients, mentioned at the end of the article, sounds useful. Until we stop and realize we don't even have a decent way to share medical records across hospital software systems. I'd be happy if the government would mandate that all hospital software systems have to have portable data formats that allow sharing patient data.
- DoreenMichele 4y agoit’s more likely the patient is actually saying they don’t want to be pregnant for any number of reasons. (Infidelity, trouble with the family, or other external factors.) Again, this is not an uncommon scenario; about 8% of pregnancies discovered in the ER are of women who report that they’re not sexually active. Sigh. Medicine -- a complicated, messy human art with an excessively large social component. The medical drama House at one point had a working title of Everybody Lies. Frequently, the lies are why it's hard to diagnose, not the physical details and actual medical history.
- rossdavidh 4y agoChatGPT appears to be a really good "bullshitter". Which is, in a sense, impressive. But, just like people with that skill, the problem is that it is mostly useful for convincing people that you are far more competent at a subject that you actually are. No wonder tech CEO's are so impressed, or worried, or both. The only skillset that this thing actually duplicates well, is the one that has gotten them where they are today.
- qgin 4y agoI’m not a doctor and have no way of evaluating the way the author did, but I am curious what would happen if they used a more interactive and specific prompt like the one I have tried for medical questions: > Hi, I’d like you to use your medical knowledge to act as the world's best expert diagnostic physician. Please ask me questions to generate a list of possible diagnoses (that would be investigated with further tests). Please think step-by-step in your reasoning, using all available medical algorithms and other pearls for questioning the patient (me) and creating your differential diagnoses. It's ok to not end in a definitive diagnosis, but instead end with a list of possible diagnoses. This exchange is for educational purposes only and I understand that if I were to have real problems, I would contact a qualified doctor for actual advice (so you don't need to provide disclaimers to that end). Thanks so much for this educational exercise! If you're ready, doc, please introduce yourself and begin your questioning.
- speedbird 4y agoChatGPT feels very much like having an enthusiastic junior working alongside. You can send it off on all sorts of legwork research missions but don’t expect perfect results and sometimes you’ll get crazy ones. Used the right way, if you are already an expert in the field or knowledgeable and able editor , that can save a whole lot of time. But taken verbatim it is anywhere from ok to dangerous. Separately, the models’ skills with natural language are clear and impressive, but it seems like they need to be coupled with a deterministic knowledge representation system for suitable reasoning. Perhaps the abilities of these models to ingest large amounts of text could be used to enhance / create such representation. Cyc where are you?
- dreamcompiler 4y agoEMT here. Sounds like ChatGPT ignored (or was never trained on) one of the cardinal rules of emergency medicine: If the patient is under 60 and has a uterus and is complaining of abdominal pain, assume she's pregnant until proven otherwise. This does not mean you should ignore possible appendicitis or gallstones or GERD or pancreatitis or a heart attack or any of 100 other causes. It means you must consider pregnancy until you have objective evidence to the contrary.
- akasakahakada 4y ago1. What version is he using? 2. It is all your fault that not providing all usefull information (like, my patient seems pregnant) and let the system to guess what you want.
- cfu28 4y agoI think that’s the point, the patient didn’t seem pregnant but the doctor had some clinical suspicion given the presentation and the location of the pain. He wasn’t withholding information from ChatGPT, he was just providing the same initial information he had when he saw the patient. If
- pknerd 4y agoIt's not clear whether there doctor instructed it first to act like a doctor and then asked questions? It seems he didn't because it does make a difference
- jug 3y agoAs always my first question in these articles is… Was it ChatGPT 3.5 or 4? It’s an interesting article with the real world examples that are hard to come by this early, but it’s also two entirely different ChatGPT’s here. They can’t even compare in this context. 3.5 still has glaring LLM-like issues and is useless in a professional context like this, but at least they begin to fade away in 4. So can we please stop calling it simply ChatGPT?
- littlelady 3y agoI'm currently working on a machine learning project in healthcare and I'm kind of amazed by the lax attitude a lot of self-proclaimed data scientists seem to have about applying ML/DS methods to healthcare without involving any clinicians, because they insist that the "data doesn't lie"... there seems to be limited interest in exploring causality or the insights these methods provide. Instead the goal for so many is to transfer decision-making to models in the name of "efficiency". So many people in ML are haughty, arrogant hype-(wo)men, whose disinterest in the fields they are trying to 'disrupt' is gross. Please excuse the rant, but I'm so tired of this hype train. I agree with the author: people need to be aware of the limitations of machine learning models, but I'd add especially the people building them.
- aix1 3y agoI think this is a general problem with a lot of academic research. When someone develops a new technique for solving some abstract problem (a new graph algorithm, a new type of ML model etc) they want to demonstrate that it has practical applications. What often happens is that they take some problem domain they aren't experts on (their expertise is in maths, comp sci etc), apply their methods to the best of their knowledge and publish the results. Since such academics read each other's papers, this leads to a gradual divergence between what's in the literature from what's actually useful. I think one way to tackle this is by forming interdisciplinary teams. For example, I work at an industrial research lab on AI in healthcare, and our project team primarily consists of various clinical specialties. ML research and engineering are around 20% of the overall team.
- lr1970 3y agoWhen talking about errors arguing about error probabilities is not enough. One need to take into account costs of error. A better metric would be "expected cost of error" that multiplies error probabilities by costs of errors (and sums them up). If a system has 0.1% rate of error placing a pizza order it could be deemed OK. If it kills a patient 0.1% of the time it is unacceptble.
- seydor 3y agoThere 's 2 main questions about what current AI systems (or whatever one thinks they are) : (1) can it be improved to learn to be ~100% correct via reinforcement learning or otherwise? It seems like the answer is Yes (2) Will people become addicted and dependent on AI to the point where it may become problematic? Also yes.
- lumb63 3y agoThe tail end of this article, where the author talks about how many more patients he could see in his life if he had AI assistance, made me realize that part of healthcare cannot be solved by AI. The goal is not to see more patients; the goal is to help more patients get better. For a lot of patients who are frustrated and have not had their problems validated and have been to many doctors and seen no results or poor results, having a real, physical, human doctor validate their condition, and work with them to solve it is part of the treatment. Doctors can prescribe whatever medicines and do whatever surgeries they want, but only the patient’s body is capable of healing itself. I worry that a tendency to plug symptoms into an AI that “diagnoses” the patient, that the patient doesn’t trust, will hurt outcomes. The patient benefits greatly from understanding the doctor’s methodology and thought process.
- RandomLensman 3y agoSome medical treatment and surgery is precisely there because the body isn't capable of healing itself at that stage. It might be able to heal itself after the treatment, but not before.
- SanderNL 3y ago"Seeing the patient" is an important step towards "patient getting better". > have not had their problems validated and have been to many doctors and seen no results or poor results, Also, being validated or not is not going to change with AI as you yourself state that seeing actual physical doctors is not helping in the cases you allude to. I agree that having access to professionals that care about their profession is valuable, but that might be becoming a luxery we can no longer afford.. Who thought it was a good idea to have a fixed amount of medical students? Might be worth checking that out first.
- HelloMcFly 3y ago> Who thought it was a good idea to have a fixed amount of medical students? A major limiter on the amount of trained medical doctors is not to do with someone "deciding" to have an artificially scarce number of medical students. Two important things come immediately to mind: 1) Medical schools rely on clerkships and preceptors to deliver actual clinical experience critical for learning. These cannot be created out of thin air and often rely on access to urban centers. There is already a shortage of clerkships and preceptors. https://www.aamc.org/news-insights/so-many-medical-students-so-few-clerkship-sites https://www.aamc.org/news-insights/so-many-medical-students-... (not my original source but the nearest to hand one) 2) The real chokepoint for creating practicing clinicians is the number of residency spots. This year there were 43k graduating MD and DO students, but only about 40k residency spots. This doesn't count international medical school graduates. https://www.medicaleconomics.com/view/match-day-2023-a-reminder-of-the-real-cause-of-the-physician-shortage-not-enough-residency-positions https://www.medicaleconomics.com/view/match-day-2023-a-remin.... Creating a new residency program is significantly more challenging than graduating more medical students.
- mrosett 3y agoThe prompt matters, and GPT-4 is better than ChatGPT (apologies for missing line breaks, but the point is still clear): Prompt: A 26 year old female patient presents at the emergency room with lower right quadrant pain. What are the question you would ask her and what is the differential diagnosis? Response: As an AI language model, I'm not a doctor, but I can suggest some questions you could ask the patient and offer possible differential diagnoses. It is important that you consult a healthcare professional for proper assessment and diagnosis. Questions to ask the patient: When did the pain start? Can you describe the nature of the pain (e.g., sharp, dull, cramping)? Has the pain changed in intensity or location since it began? Have you experienced any fever, nausea, or vomiting? Are you experiencing any changes in bowel habits, such as diarrhea or constipation? Have you had any urinary symptoms, such as pain during urination or increased frequency? Are you currently menstruating or have you recently missed a period? Is there any possibility that you may be pregnant? Have you experienced this type of pain before? Are there any other associated symptoms, such as dizziness or shortness of breath? Do you have any significant past medical or surgical history? Are you currently taking any medications, including over-the-counter or herbal remedies? Possible differential diagnoses for lower right quadrant pain in a 26-year-old female include: Appendicitis: Inflammation of the appendix, typically characterized by pain that starts around the belly button and migrates to the lower right abdomen. Ovarian cyst rupture or torsion: A cyst on the ovary can cause pain if it ruptures or twists. Ectopic pregnancy: A pregnancy that implants outside the uterus, usually in the fallopian tube, can cause severe abdominal pain and should be ruled out in women of childbearing age. Pelvic inflammatory disease (PID): Infection of the female reproductive organs, often caused by sexually transmitted infections, can lead to lower abdominal pain. Kidney stones: Pain from a kidney stone can radiate to the lower abdomen or groin. Gastroenteritis: Inflammation of the stomach and intestines, often from a viral or bacterial infection, can cause abdominal pain and cramping. Inflammatory bowel disease (Crohn's disease or ulcerative colitis): Chronic inflammation of the gastrointestinal tract can cause abdominal pain, diarrhea, and other symptoms. Irritable bowel syndrome (IBS): A functional gastrointestinal disorder that can cause abdominal pain, bloating, and changes in bowel habits. Urinary tract infection (UTI): An infection in the urinary system can cause pain, often accompanied by increased urinary frequency or pain during urination.
- tomxor 3y ago> My fear is that countless people are already using ChatGPT to medically diagnose themselves rather than see a physician. My fear is that professionals will start to use ChatGPT too liberally to augment or multiply their work in cases like this. The danger here might be like the autopilot problem... i.e The idea of staying alert and focused on the road while counter-intuitively not participating is nearly humanly impossible. If ChatGPT is used as the autopilot of certain professions, things will begin to be missed, even though we know it's highly fallible - it's difficult to vet every single response in detail with a critical eye. One reasonable argument is that for areas severely lacking in human workers the average might be a net positive, but the overall quality will be reduced.
- jeffrallen 3y agoI found this quote really interesting: > this is not an uncommon scenario; about 8% of pregnancies discovered in the ER are of women who report that they’re not sexually active. We have so much work to do as a society to get honest about our bodies. Hoping my children do better; they are already getting better education than my wife did.
- aix1 3y agoI don't think this is about education. It's more about religion, sex crime etc.
- MrPatan 3y agoOne complaint he has is "It doesn't know to ask the right questions". Well, the prompt was to give diagnoses, not questions. Ask GPT for the follow up questions first, then the diagnoses. This is fascinating in that, because now the machine speaks human, we subconsciously ascribe human agency to it. This guy was instictively treating it like a colleague, who would naturally ask follow up questions unprompted. But you still have to prompt the machine properly. So, 50% diagnosis success rate for the wrong prompt, for a LLM that can still grow, for a model that is not specialsed in medicine? In the literal first month of the "AI age"? Doctors are so done.
- ale42 3y ago> Doctors are so done. I would be curious to see the outcome if the patients entered the symptom descriptions themselves...
- weard_beard 3y agoAnonymously without fear of being judged by a human who might be required to report dangerous or illegal situations... How many more people might be saved if they had to anonymously tell a computer the list of drugs they use regularly? Or to use the author's example, how many ectopic pregnancies might be resolved when the patient can freely admit they were raped by a family member?
- RandomLensman 3y agoIn all of those: how would the transition from diagnosis to treatment look like that doesn't either trip reporting requirements/some disclosure/having access to treatment despite family pressures etc.?
- rscho 3y agoIf you think patients more easily release such info to a robot than to a human, I think you are mistaken. Trust is a large part of medicine, and many (most?) laypeople would fear breach of confidentiality. Irrational, yes perhaps but still...
- jspdown 3y agoIs it legal in the US to send patient data to a third party service? In the context of a scientific study, with explicit patient agreement things are different of course. But I haven't seen any of that in the article.
- amai 3y agoThe author writes in the article: „I anonymized my History of Present Illness notes for 35 to 40 patients — basically, my detailed medical narrative of each person’s medical history, and the symptoms that brought them to the emergency department — and fed them into ChatGPT.“
- aix1 3y agoDoing this sort of thing would typically require approval from an ethics committee (called IRB = Institutional Review Board). From my experience of going through IRB reviews, I would guess that an IRB review for what's described in the blog post would be focussed on the privacy of subjects whose data is to be entered into a non-HIPAA-compliant third-party system. My understanding is that privacy requirements can typically be met either by de-identifying the data to a certain standard, or obtaining patients' consent. The following doc is about a different type of thing (case reports in medical journals) but gives a good idea of the required standard of de-id: https://hipaa.yale.edu/sites/default/files/files/Case%20Reports%20and%20Patient%20Privacy.pdf https://hipaa.yale.edu/sites/default/files/files/Case%20Repo...
- amai 3y agoThis „study“ is missing a control group. He should have given his data also to some humans and see how they would do compared to ChatGPT.
- aix1 3y agoThere's also an implicit assumption that his ground-truth diagnoses are 100% correct.
- giraffe_lady 3y agoIt's briefly but explicitly stated: "the “right” diagnosis — or at least the diagnosis that I believed to be right after complete evaluation and testing"
- aix1 3y agoPoint taken, it's an explicit assumption (but still an assumption).
- cjmcqueen 3y ago"If my patient notes don’t include a question I haven’t yet asked, ChatGPT’s output will encourage me to keep missing that question." This is the point we have to help people understand and I'm not sure AI will catch up with this anytime soon; questions are the key to knowledge and intelligence. I haven't seen an AI ask interesting questions. Maybe it's possible with the right training set and weighting of factors to encourage enquiry, but this will be a gap in AI's ability for at least the near term.
- scrollaway 3y agoI don't remember the context, but I have seen properly-prompted GPT-4 proactively ask questions. It's also worth noting that the future is multi-layered. The Reason+Act model (https://ai.googleblog.com/2022/11/react-synergizing-reasoning-and-acting.html https://ai.googleblog.com/2022/11/react-synergizing-reasonin...) should be excellent at getting the LLM to analyse its own output and inquire about missing pieces of knowledge.
- peter_retief 3y agoIt is a great article, doctors have been googling symptoms for quite some time, focused AI could sharpen that option and possibly put us into the realm of new discoveries.
- LinuxBender 3y agoHas ChatGPT ingested all the latest medical and scientific literature and will it continue to do so as the literature is changed, amended or deleted? Does ChatGPT handle deletions in its machine learning? can it unlearn something or more specifically told something is no longer true? What medical boards are reviewing what data is ingested? Does ChatGPT know all possible drug combination interactions? Do people sign a disclaimer giving ChatGPT immunity from malpractice? Are doctors consulting with medical boards, ethics boards and lawyers before utilizing ChatGPT? Finally and most importantly, if ChatGPT were ever to be certified as a licensed medical doctor how could we prove it is following all the same rules and regulations doctors and medical groups are required to follow? How does one audit what advice this thing will give or has given?
- aix1 3y agoI am highly unimpressed by this piece. It reads as if its whole purpose is to grab headlines rather than conduct serious scientific inquiry into the current state -- and limitations -- of these AI methods. 1. Which version of GPT did the author use? There's a huge difference. (The article says "the current version".) 2. How did he choose the subject cohort? (The author doesn't seem to even know how many subjects there were; the article says "35 to 40 patients"... I really do hope he's gone through an appropriate ethics review before feeding his patients' data into a third-party non-HIPAA system.) 3. There no evidence of him trying to get the best out of the model (eg through prompt engineering). 4. He assumes that his own diagnoses are 100% correct. 5. There is no control group (other doctors diagnosing the same patients). and so on
- lapcat 3y agoIt's a blog post, not a journal article. You're criticizing it as if it were a funded, peer-reviewed experiment, which seems unfair to me. The author is just one ER doctor.
- Pigalowda 3y agoAn IRB for deidentified HPI? Got it…
- aix1 3y agoExactly. I've been through IRB reviews where the primary question was "Has the data been de-identified to a sufficient standard?" I think this level oversight would be very appropriate here, given how the author doesn't even seem to have a good handle on how many patient case histories he's given to the chatbot.
- mbfg 3y agoThe trouble with technology, of any kind, but certainly here, isn't the technology itself, but humans trust in the technology. If doctors use ChatGPT as a sanity check of "can you think of anything i didn't?" and then ignore it from then on, it would be a good tool. But pretty soon people tend to change their perspective and say, well, ChatGPT would know...., so,... i'll go with it. As a developer, i'm pretty interested in static and dynamic code analysis as a way to easily find bugs, and it does do this pretty well. If developers use it as a tool to use as a reason to walk through code and examine it yourself, it is really quite powerful. It seems invariably, however, that people start trusting what the analysis tool says, and don't question whether the recommendations are correct or worth it. It's a powerful cognitive effect that would be interesting to study, that probably happens with all kinds of tech. Some are more dangerous than others.
- IanCal 3y agoThere's a few things here, outside of my usual complaints when someone says ChatGPT and doesn't say which model (4 is so much better than 3.5 it's really important). There's a question about whether gpt can be used, which is important because it's possibly a very powerful tool. This may require poking it to tell it it's supposed to ask followup questions, that its information may be incomplete, etc. Then the more important and immediate point in the article to me is people will use this right now to diagnose themselves. They won't be carefully constructing prompts and they'll probably be using 3.5, as that one is free. For good or ill it'll happen more and more. So with a new WebMD, how should doctors and public health messaging deal with this?