4 ms·
That's wildly misleading then. It would be interesting to see how GPT-4, properly augmented with actual medical literature, would do.
by CSMastermind 3y ago
That's wildly misleading then. It would be interesting to see how GPT-4, properly augmented with actual medical literature, would do.
- AustinDev 3y agoLikely better than the average doctor. If I had the opportunity to take that bet, I would.
- aurareturn 3y agoI agree. While I appreciate what doctors do, there sure are a lot of shitty doctors out there who skirts by - like any profession.
- intended 3y agoI would take the other side of that bet in a heart beat. But given the vagueness of the wording, much is going to depend on the details.
- RobinL 3y agoI wonder if th result changes if you put a high quality medical reference in context. Feels like there might be an opportunity for someone to try and cram as much medical knowledge as possible in 1m tokens and use the new Gemini model.
- threecheese 3y agoSame; a doctor’s judgment is supported by a system of accountability, which distributes the risk of error beyond the patient to the doctor/medical practice/insure. In contrast, (at least as of today) user facing AI deployments absolve themselves of responsibility with a ToS. Who knows if that’ll stand up to legal scrutiny, but if I have to bet on something ITT it would be that legal repercussions of bad AI will look a lot like modern class action lawsuits. I look forward to my free year of “AI Error Monitoring by Equifax”.
- catwell 3y agoI saw a presentation about this last week at the Generative AI Paris meetup, by the team building the next generation of https://vidal.fr/ https://vidal.fr/, the reference for medical data in French-speaking countries. It used to be a paper dictionary and exists since 1914. They focus on the more specific problem of preventing drug misuse (checking interactions w/ other drugs and diseases, pathologies, etc). They use GPT-4 + RAG with qdrant and return the exact source of the information highlighted in the data. They are expanding their test set - they use real questions asked by GPs - but currently they have 0 % error rate (and less than 20 % cases where the model cannot answer).