4 ms·
I recently published a paper, where we explain how an FDA approved prediction model, build into a widely used cardiac monitor was developed with an incredibly b
by johannes_ne 4y ago
I recently published a paper, where we explain how an FDA approved prediction model, build into a widely used cardiac monitor was developed with an incredibly biased method.
https://doi.org/10.1097/ALN.0000000000004320 https://doi.org/10.1097/ALN.0000000000004320
Basically, the training and validation data was engineered so an important range for one of the predictor variables was only present in one of the outcomes, making perfect prediction possible for these cases.
I summarize the paper in this Twitter thread: https://twitter.com/JohsEnevoldsen/status/1561641153899929601?t=jVcYm9J9wGT1F-Y_LpC9lQ&s=19 https://twitter.com/JohsEnevoldsen/status/156164115389992960...
- baxtr 4y agoSorry for asking, but how is this relevant to the article?
- NovemberWhiskey 4y agoSorry for asking, but how is it not?
- baxtr 4y agoDo you agree that it’s ok to pose a question whenever you don’t understand?
- csallen 4y agoIronically, that's exactly what NovemberWhiskey is doing here :)
- ShamelessC 4y agoI’m not sure where you got this form of communication where you respond to everything with a question, and I assume you mean well, but it comes across as patronizing and de-humanizing to try to follow these “rules to winning arguments passively”, or whatever it is. Indeed, the confusion here is (I think) because your first comment > Sorry for asking, but how is this relevant to the article? Sounds accusatory. Please don’t respond to this with a question.
- baxtr 4y agoThe basic idea of that kind of question is to find the minimal place of agreement. And then understand where one deviates. Going back the path of arguments to common ground if you will. It works quite well in my experience if you’re interested in genuine discussion. PS: how something “sounds” is really difficult to say in a written medium. It might say more about the reader than the writer.
- isitmadeofglass 4y ago> PS: how something “sounds” is really difficult to say in a written medium. It might say more about the reader than the writer. No, it’s not difficult. And not it’s not the reader. When multiple readers all agree about the same interpretation of the writer. It might have been unintentional on the part of the writer, but that doesn’t make it “difficult” Or the readers fault.
- ShamelessC 4y agoI think what youre trying to encourage is open ended discussion? It's my opinion that this only tends to work IRL or in online mediums with more moderation e.g. wikipedia, stackoverflow. Random open ended discussion can be good, but I bet it's wise to assume tht most random musings arent really as interesting as you might think. In any case thanks for clarifying.
- johannes_ne 4y agoFair question. The model we comment on both suffer from the problem described in the article but also a more severe problem: The developers sampled obvious cases og hypotension and nonhypotension, and trained the model to distinguish those. And also validated it on data that was similarly dichotomous. In reality the outcome is often between these two scenarios. But worse, they also introduce a more severe problem where as range of an important predictor is only available in the hypotension outcome.
- baxtr 4y agoThanks for explaining!
- mjthrowaway1 4y agoI quit research forever after I was ignored pointing out a similar problem in our predictive model.
- johannes_ne 4y agoI can only imagine the frustration. Just getting this through peer-review took half a year, but at least there was the academic currency of a publication to motivate me.
- roflyear 4y agoThis has to be intentional no?
- 77pt77 4y ago> I decline to answer that question on the grounds that it might be used to incriminate me
- johannes_ne 4y agoThe problem is quite subtle, though obvious in retrospect. I've seen a paper from a separate, academic, research group make similar model with the exact same problem. The problem would, however, have been clear, if the model was compared to simply using the current mean blood pressure (MAP) as a predictor of hypotension, because MAP is the problematic predictor variable. Instead, the model was only compared to short-term changes in MAP (ΔMAP), which is obviously nonsensical and has an AUROC of ~0.55.
- pas 4y agoHm, reading the linked tweets the problem seems like a big screaming red target on the side of a white barn, not a feature engineering subtlety. It seems like the typical case of the drunk guy looking for his keys under the streetlight. (Having insufficient data, and comparing the model to an arbitrarily picked one that just happens to be even worse. And then everyone including the FDA patting them on the back.)
- johannes_ne 4y agoI'm glad that you seem to get the severity! I'm just hesitant to ascribe malice.
- pas 4y agoI think it is the general incompetence of the "academia + R&D biz + regulation pipeline". (In the land of the blind the one-eyed is king, etc.) It's sort of inevitable in such a non-teleological process. As in each step in it serves its own purpose, and so the whole thing doesn't really serve the purpose that we like to assume for it - ie. give us great thoughtful inventions. That's why it took so long to stop the Theranos train, that's why it takes so fucking long to roll out polyvalent vaccines (ie. all-in-one vaccines), and so on. (I'm picking on medtech here but there are many others, the Boeing + FAA MCAS fuckup, the absolute limpdick paralysis of nuclear power - it needed a combination of half the world on fire + prelude-to-WWIII to get it moving again, and so on.)