3 ms·
> Reminder once again that "accuracy" is irrelevant in the real-world. Accuracy and other quantitative performance metrics are imperfect for sure, but how else
by EForEndeavour 2y ago
> Reminder once again that "accuracy" is irrelevant in the real-world.
Accuracy and other quantitative performance metrics are imperfect for sure, but how else do you propose testing before real-world deployment? How do you propose scalable and feasible testing of human students?
> The practice of medicine does not produce cute little cue-card prompts with 4 options.
Ah, but the New England Journal of Medicine's Image Challenges (designed to test the knowledge and diagnostic capabilities of medical professionals) does.
> It's of less-than-zero value to have an 80% machine accuracy and 75% clinician accuracy if the impact of the machine's mistakes are high risk, unconcerned with the impact of the intervention on the patient, or provide little therauptic upside
But this paper does not study the practice of medicine. It intentionally focuses on performance on one specific, well-known medical imaging diagnostic challenge.
- mjburgess 2y agoDefine a utility across the outcome distribution, eg., profit/quality-adj-life-years-etc. 95% of ML "research" would reveal itself pretty useless if people did this, since the "80-90%" accuracy we're getting is on the "break-even" part of the profit-curve. We're not getting innovation. It's very rare that frequentist stat modelling on historical data would produce anything, in this sense, suprising -- ie., suprise in the utility domain