9 ms·
I have a friend who is an OB/GYN Oncologist in Indianapolis. Her hospital had IBM Watson Health on campus last year, and she mentioned to me in passing last mon
by imroot 7y ago
I have a friend who is an OB/GYN Oncologist in Indianapolis. Her hospital had IBM Watson Health on campus last year, and she mentioned to me in passing last month that they finally had kicked IBM off of their (learning) university's campus. When I asked her why, she said:
"Often times, Watson would recommend courses of treatment that would be completely incorrect, if not detrimental -- or even sometimes lethal -- to the patient. It became more of a hassle than a learning tool."
- devoply 7y agoIt's good we still have doctors to provide oversight for this sort of thing. Imagine a future where machines take over because they are thought to be smarter than doctors and then start killing people and their mistakes are written off because the people were already sick. Like an Elizabeth Holmes AI.
- euler_ 7y agoDoes it make a difference if they preform better than human docs? The same thing happens now with human docs.
- ethbro 7y agoThe difference is the expectation of fallibility with human docs. Which is why checks are built into the system. I have zero trust in IBM to market their ML products correctly so that proper checks are maintained. Especially since technically, explainability is still an active area of ML research.
- freehunter 7y ago>I have zero trust in IBM to market their ML products correctly so that proper checks are maintained. That's a problem, yes, but it's not a new one. Vendor management has been around for decades, centuries, millennia maybe. If I contract out part of my job, it's still my responsibility to make sure the contractors are doing their job right. "But their marketing said..." or "but their sales guys said..." is not an excuse and everyone knows that. Doctors noticing that Watson is wrong is expected. Doctors missing the fact that Waston is wrong is a failure of that doctor and the doctor who didn't check the results is the responsible party. The checks don't come from Watson, the checks come from humans who oversee Watson. If Watson is wrong often enough that it's hindering the doctors, then kicking it out is the right call. But there can never be an argument of "Watson got the diagnosis wrong and that's why the patient died" because ultimately IBM is still just a vendor and Watson is still just a contractor.
- ethbro 7y agoThe difference is like:like vs like:unlike, and seems to be one of the more dangerous ML application challenges. If I as a medical provider hire a remote vendor, who has medical teams in India look over initial results to flag issues, those humans will fail in human ways. I can anticipate that: I'm a human. If I use a similar ML product, it's very difficult for me to anticipate (or even understand) the ways it which it might / does fail. Which makes it unlike my previous experience. Which gives it a fundamentally different risk profile. It's the Boeing issue in a nutshell: the failure case that unfolded was unlike the scenarios the pilots were trained for. Unfortunately, in the two crashes they were unable to dynamically RCA quickly enough to solve the problem. My point was that coupled with IBM's inept and inaccurate marketing, it seems unlikely the appropriate risk information is in the hands of those responsible for managing risk. And honestly, if a system has unlimited failure modes, and I can't learn and limit them in practice, it's useless. Because in that scenario I should be duplicating all the work it's done to ensure it didn't go off the rails. In practice (and guided by labor cost savings promised in the contract signed with management), that full verification doesn't happen (because the vendor is incentivized to recommend it doesn't), and people die.
- freehunter 7y agoIt's a good point that it's a failure of management. It almost always is in situations like this. I design and deploy customized automation systems for customers, and it's part of the standard process that we run the automation side-by-side with the old process for several months in order to learn the new failure methods and synchronize the process. Yes, for a few months we're duplicating the machine's work, but without the machine we'd be doing the work anyway. And no one is going to die if my automation fails, but we still do this anyway. It's crazy to think anyone would believe they didn't need to do side-by-side verification no matter what sales and marketing told them. I don't know enough about Watson or IBM sales to say if Watson is good or bad, but I'm not trying to defend Watson or IBM. Watson may very well be a complete failure. But that aside, it's not the only failure in this story. No one should expect to implement a new tool and never verify if it's working correctly.
- tellak 7y ago
- Bartweiss 7y ago> the expectation of fallibility with human docs Not just the expectation but the understanding. A doctor might very well forget which leg to amputate, so we know to Sharpie "NOT THIS LEG" on the one being kept. But a doctor is very unlikely to see a patient with a broken wrist and prescribe antipsychotics, so we don't do much to prevent that error. Human fallibility happens along fairly predictable channels, and we've spent a very long time committing resources to controlling those channels. Watson, though, thought Toronto is a city in the USA. Anyone who's dealt with ML output knows that the errors are often quite surprising, even before dealing with adversarial inputs. Even in a system where Watson's outputs are subject to checks, the checks we have today are human-specific and developed at a significant human cost. ML answers can't just outperform individual human doctors to add value, they need to either be gracefully integrated with them or be able to outperform the entire system which keeps those doctors on track.
- tellak 7y agoI don’t understand this argument (which comes up often). Isn’t the answer an unqualified yes? It’s obvious the current system has fuck ups. Are we supposed to be happy if a new system has double or triple the fuckups because “technology”? Like where are you going with this? These clinical decision systems have not proven they improve outcomes and we have no reason to believe they won’t make things worse.
- TheOtherHobbes 7y ago"We successfully automated lethal incompetence" is perhaps not the most appealing of all possible USPs for a medical system.
- Waterluvian 7y agoThis reminds me of a bit in the first The Expanse book I read recently. The onboard medical computer had to be overridden because it calculated coldly that the correct course of treatment was palliative care.
- plttn 7y agoGiven the context of that scene, it was more an indication of how much characters A and B had been affected by the situation.
- Waterluvian 7y agoYeah. Not trying to draw a parallel. Just reminded me of sci fi.
- MartinCron 7y ago...in the TV adaptation anyway, it is one of the best bits of dark humor.
- tabtab 7y agoThe purpose of such a system should be to give leads for further research. If you treat them as leads and only leads, then bad leads don't matter much. Unless, the leads are so bad on average that too much effort is spent on vetting them that could be put to better use elsewhere.
- arkitaip 7y agoI doubt that's how Watson is being marketed to customers, though. IBM tends to oversell and under deliver and have grown increasingly desperate with their shrinking market shares in both hardware and services.
- coldtea 7y ago>The purpose of such a system should be to give leads for further research. If you treat them as leads and only leads, then bad leads don't matter much. If the leads are worse than a random toss of a coin, then you can come up with leads on your own. Not all leads are good. Inviting the village idiot in a brainstorming session wont be of much value -- and Watson is more like the village idiot, than a valid lead generation engine.
- skwb 7y agoThese are only a few data points in a almost half decade of promising for healthcare solutions and failing to meet those expectations. The MD Anderson audit is particularly bad: https://www.utsystem.edu/sites/default/files/documents/UT%20System%20Administration%20Special%20Review%20of%20Procurement%20Procedures%20Related%20to%20UTMDACC%20Oncology%20Expert%20Advisor%20Project/ut-system-administration-special-review-procurement-procedures-related-utmdacc-oncology-expert-advis.pdf https://www.utsystem.edu/sites/default/files/documents/UT%20...
- save_ferris 7y agoWow, they didn't waste any time in that executive summary. Those bullet points are insane on their own.
- dmix 7y agoEven ignoring all the lack of deliverables, just the amount of financial trickery and lack of project oversight by the way money was pumped into it via philanthropy (instead of primary funding channels) which meant it bypassed multiple checks and balances, including from anyone knowledgeable of software development. Sounds like every other failed $50+ million government software project ever, which IBM is apparently an expert at. Except this money donated by private citizens was redirected from other cancer research projects and siphoned into a billion dollar company’s coffers. Really shameful stuff.
- burtonator 7y agoReminds me of this: https://techcrunch.com/2009/09/02/netbase-thinks-you-can-get-rid-of-jews-with-alcohol-and-salt/ https://techcrunch.com/2009/09/02/netbase-thinks-you-can-get... ... granted this happened a decade ago: > "Several of our readers tested out the site and found that healthBase’s semantic search engine has some major glitches (see the comments). One of the most unfortunate examples is when you type in a search for “AIDS,” one of the listed causes of the disease is “Jew.”