4 ms·
No method is fool-proof for GPT-3, but there are tricks that suppress its tendency to hallucinate, such as demonstrating in a k-shot prompt the possibility of f
by goodside 4y ago
No method is fool-proof for GPT-3, but there are tricks that suppress its tendency to hallucinate, such as demonstrating in a k-shot prompt the possibility of false assumptions in a question. I demonstrate this here: https://mobile.twitter.com/goodside/status/1545793388871651330 https://mobile.twitter.com/goodside/status/15457933888716513...
Models more advanced that GPT-3, such as LaMDA, have entire subsystems specifically dedicated to “grounding” output in truthful information. Hallucination is at least partially a solved problem, but the methods haven’t disseminated broadly yet.
- abrax3141 4y agoUnless it’s reading and (in some sense) understand FDA and CDC drug approvals, ct.gov, pubmed, and NCCN and other treatment guidelines, and realizing when drugs are suddenly deprecated as well as suddenly approved, it’s going to commit mal practice. You could do all that, but then you don’t have a language model, you have a domain specific medical AI.
- goodside 4y agoNot sure where medicine even entered the discussion. Obviously GPT-3 is not a doctor and shouldn’t be used as one. That’s a much harder problem than suppressing the tendency of GPT-3 to confabulate/hallucinate fictional memories.
- abrax3141 4y agoIt’s not obvious. Medical advice is among the top searches on google - possibly the top. And is, by that very same virtue, among the first things people are going to ask what they hysterically believe to be an AI.
- goodside 4y agoThere’s a big leap from “some layperson might misuse this for medical advice” to “code generated by GPT-3, for any purpose, cannot be reviewed and fixed by expert human programmers,” which is where this thread started. I’m not sure what you’re arguing, or if you’re even trying to support that original point.
- abrax3141 4y agoWe should probably cut this off, but let me try to explain what I had in mind, which might be none sense, but was certainly not well explained. Too briefly, presumably we can agree that it’s even hard for experts to check and fix code created by other humans effectively. Now put this in a dangerous domain. I originally said aviation, but (confusingly) switched to medicine for reasons not worth going into, but we can replace “program” with “treatment regimen” and “expert programmer” with “specialist physician”. The fail I was reaching for is where what is hallucinated by the LLM looks really convincing (has great grammar), and is close enough to correct that it isn’t glaringly wrong - and these models are good at doing those two things! - but by virtue of the fact that they can’t reason, they aren’t actually “thinking things through” - which is why, of course, you need the programmer/specialist. But the hallucination is so good (yet wrong, per unreasoned) that users (patients) are taking the word of the model. You think that can’t happen, but we do it all the time with results from “Dr Google”. A friend who is a doctor says that patients regularly come in with inch think binders of stuff they printed off the web that he has to talk them down from. We’re headed for a near future where instead of being talked down from stuff you merely printed off, you’ll have to be talked down from an extremely convincing wrong treatment hallucinated by Dr HAL. Anyway…never mind.
- kreetx 4y agoI think "put this in a dangerous domain" was what the parent was specifically not doing. I also wonder if it's GPT-3 itself arguing as the text is kind of strange. :)