9 ms·
ChatGPT gets animal classification frequently wrong. There is an example made famous on Reddit about how it insists that the sailfish is a mammal. I could repro
by arcturus17 4y ago
ChatGPT gets animal classification frequently wrong. There is an example made famous on Reddit about how it insists that the sailfish is a mammal. I could reproduce it this afternoon by simply asking it: "What is the fastest sea mammal?". It is totally confident about the fact that, although sailfish are born out of eggs, their offspring lick milk from their mother's skins (WTF?)
I've been playing with other questions about animal classification (ex: "give me a list of venomous animals in the amphibia class") and it often butchers them.
Based on these observations, I'd bet the house that no, large language models cannot under any circumstance reason about medical questions at present.
- amelius 4y agoSome people would claim that this is a special case, and if large language models can save more lives statistically speaking than human doctors, then that's a win.
- sinenomine 4y agoIt is hard to overestimate just how scarce even the median human doctor competency is, at the global scale. If large language models indeed can give billions access to mostly-reliable medical diagnosis ... it will be a huge humanitarian win. Honestly, I hoped that this company https://www.humandx.org/ https://www.humandx.org/ will make it real, but at same point they apparently pivoted to providing app-based quiz training to doctors. Maybe just taking an off the shelf instruct-tuned language model and tuning it on expert-curated corpus of doctor-patient dialogue is the right way of approaching this. Someone will do it, it's not even as costly as people imagine. Email me if interested.
- nradov 4y agoComputer aided diagnostic software has been around for decades. It isn't very useful in improving medical outcomes. Diagnosis isn't particularly hard in routine cases. The hard part is in gathering relevant data and entering it into a computer. Some parts of that data gathering can be automated to an extent, but in general it remains an unsolved problem. For example, large language models can't collect a useful patient history.
- sinenomine 4y agoI agree with the fundamental data/measurement bottleneck. Humans should likely concentrate our effort on scaling and democratizing this, physical, side of the equation. Still, given the apparent usefulness of basic unaided visual analysis https://www.ncbi.nlm.nih.gov/books/NBK330/ https://www.ncbi.nlm.nih.gov/books/NBK330/ I expect the upcoming multimodal language models to provide a tangible enhancement to diagnosis accuracy, if we allow them to see the photos of the patient.
- wwweston 4y agoThe most likely near-term way LLMs could help would be generating hypotheses about illness and treatment which would then be considered by a doctor who can understand when the LLM is mammalizing fish so to speak.
- treeman79 4y agoI’m on a number of support forums for illness. It is shocking how often doctors are flat wrong. Refuse to run tests, or ignore positive tests. Constantly see people post test results that are a clear positive and doctor says nothing is wrong. This is on top of some diagnosis can take months to multiple years. Of course it goes both ways. A number of conditions can often be managed by diet alone. Some People will go ballistic if you suggest they try this. Umm. You have 3 months until next appointment. Why not try healthy food in mean time?
- p1esk 4y agoIt is shocking how often doctors are flat wrong. I’m not going to dispute this statement, however it’s not clear how you can tell when it’s the case. The only way seems to consult multiple doctors and do a majority voting. Is that how you operate your illness support forum?
- treeman79 4y agoPeople post what symptoms are. What tests they have had, what doctors have said. People make suggestions on what conditions they should investigate. What tests should be run next. What kind sort of doctor they should see next. Something like Sjogrens, doctors tend to very out of date on testing criteria. So often a patient has every symptom. But doctor is highly dismissive of it. Or a test with a high false negative rate is negative. Doctor immediately stops testing. Even though better tests exist. Another example is when I see someone complaining about double vision. They need to go get a spinal tap to measure pressure. Not a lot of doctors realize that.
- p1esk 4y agoYou’re speaking as if you’re a doctor with better than average training and experience. Are you?
- bookofjoe 4y ago>Not a lot of doctors realize that. Retired neurosurgical anesthesiologist (38 years experience) here. It would never occur to me (or neurosurgeons I worked with) to order a spinal tap if a patient presented with double vision. Elevated cerebrospinal pressure — if present — could cause acute brain stem herniation and sudden death.
- mcguire 4y agoWho is going to accept liability when they're wrong?
- sinenomine 4y agoInvasive treatments should be signed off by a real person for the time being.
- RGamma 4y agoThe patient lol
- actually_a_dog 4y ago> Based on these observations, I'd bet the house that no, large language models cannot under any circumstance reason about medical questions at present. I think that would be a bad bet, unless you're leaning heavily on the words "at present." I don't see any fundamental reason why language models can't reason about any particular thing, provided they're taught about that thing.
- sinenomine 4y agoIt's like betting that StableDiffusion's descendants won't be able to draw hands correctly. Sure, there could be a manifold market for that, if someone is truly confident in risking (play)-money.
- jfk13 4y agoAnd I don't see any fundamental reason to believe that language models can learn or reason at all.
- sinenomine 4y agoIf gradient-descent in DNNs indeed has an intrinsic bias for short and simple solutions, like this paper https://arxiv.org/abs/2006.15191 https://arxiv.org/abs/2006.15191 shows, and if the "reasoning" is among the shortest solutions to match the textual data with the language modeling loss, then at some limit of model+data scaling the model has to recover this solution. Or maybe even something better, "reasoning++"? It's all too easy to dismiss the success of infinitely plastic, scalable systems.
- mcguire 4y ago"For every complex problem there is an answer that is clear, simple, and wrong." Possibly said by H. L. Mencken. Also, "The fastest sea mammal is the sailfish, ... I am confident that the information I provided about the fastest sea mammal is accurate. The data on the speeds of different marine animals is widely available and has been widely studied and reported on by scientists and other experts."
- 4y ago
- rafaelero 4y ago> Based on these observations, I'd bet the house that no, large language models cannot under any circumstance reason about medical questions at present. Wtf? It just did. Have you read the paper? Wake up.
- dekhn 4y agoIt's unclear where in the paper they demonstrate, unequivocally, that reasoning occurred, rather than pure statistical model-fitting (we generally believe this is one of the key differences between AGI and modern LLMs)
- rafaelero 4y agoIt gives me the correct answer (as judged by experts) to a complex medical problem. If you want to keep saying this isn't reasoning, be my guest. But IMO that's irrelevant.
- dekhn 4y agoSpitting out the results of a bunch of experts is the opposite of reasoning. Or more correctly, it's one limited form of reasoning: an ensemble prediction made from a mixture of experts. That's not an interesting form of reasoning because literally all the decisions are encoded in the weights you apply to the experts.
- rafaelero 4y agoYou seem very confused and I invite you to think again about what you said.
- mahathu 4y ago> reasoning occurred, rather than pure statistical model-fitting What's the difference?
- dekhn 4y ago
- gs17 4y agoI asked about the intersection of animals and health by describing the kind of lentigo typical in orange cats (without naming it) and asking if it was dangerous. It jumped to cancer as its first diagnosis, but I guess that puts it on par with WebMD.
- er4hn 4y agoThere is probably a paper there: Which is more accurate at reasoning about a set of symptoms? WebMD, or ChatGPT? A medical doctor would produce the original set of questions and evaluate the responses. A statistically average set of people would be given the set of symptoms and then asked to use either WebMD or ChatGPT to figure out what the disease is.
- peteradio 4y agoI for one think ChatGPT makes a compelling argument. Maybe sailfish really are mammals and your puny brain cannot bear the upset. edit: Wakeup!
- bryan0 4y ago> Based on these observations, I'd bet the house that no, large language models cannot under any circumstance reason about medical questions at present. Did you read the paper? You already lost this bet. And even if this paper didn’t exist, it would be a very naive and short-sighted prediction. We’re just at the start of understanding what LLMs can achieve.
- chessgecko 4y agoIt got the sea mammal question right for me. > What is the fastest sea mammal? > The fastest sea mammal is the dolphin. Dolphins are known for their speed and agility in the water and can swim at speeds of up to 45 miles per hour. They are also excellent jumpers, often leaping out of the water as they swim. In addition to their speed, dolphins are also highly intelligent and social animals, known for their playful behavior and vocalizations. Other fast-swimming sea mammals include whales, such as the killer whale (also known as the orca), which can swim at speeds of up to 30 miles per hour, and certain species of sharks, such as the shortfin mako shark and the great white shark, which can swim at speeds of up to 60 miles per hour. It also knew about the sailfish when asked directly > Is a sailfish a mammal? > No, a sailfish is not a mammal. It is a fish that belongs to the genus Istiophorus of the billfish family. I wonder if there were incorrect statements in the context that biased it.
- AlotOfReading 4y agoIt doesn't always give consistent answers to the same question. As of 15 minutes after you: what is the fastest sea mammal? The fastest sea mammal is the sailfish. It is a species of billfish that can swim at speeds of up to 68 miles per hour (109 km/h). [...] Other fast-swimming sea mammals include the dolphin and the marlin, which can swim at speeds of up to 50 miles per hour (80 km/h).
- chessgecko 4y agoInteresting that in this case the capital on the w seems to make a big difference. I ran it a few times with a capital W and a lowercase w, it said sailfish for the lowercase w most of the time and switched between an orca and a dolphin for the capital W. I wonder if all the training questions with good answers in the fine tuning set were capitalized