5 ms·
For one thing, no matching was done (as the email explains).
by cvinn 5y ago
For one thing, no matching was done (as the email explains).
- bsdetector 5y ago"if they said, we have 8 people with COPD in the first group and then waited for COPD patients to show up at the ER for the second group until they had 8" Did you not read this? I asked whether if they did this they would come up with the perfect 1 scores on those items. Whether they actually did or not is a separate matter. What troubles me more than whether this study was faked or not is the certainty HN readers have that it was without being able to answer simple questions about why they believe that. Is this test just measuring the likelihood of each group having the same primary diagnosis? If they added people to the second group so they had the same number of a primary diagnosis, would that result in a 1.0 p value on this test? These should not be difficult questions to answer for somebody certain that fraud occurred.
- civilized 5y ago1. The procedure you describe would make patient recruitment non-consecutive, contrary to the reported procedure. Consecutive means you don't skip anyone who has the condition you're trying to treat. You would have to skip people if you're selecting matched treatment and control groups. In this case, people with sepsis who don't fit your matching design have to be skipped and excluded from the study. 2. Patient subgroup counts matched perfectly not only on COPD, but on about a dozen other conditions. Difficulty of matching subgroup counts grows rapidly with number of dimensions. To match subgroup counts for Group A and Group B near-perfectly on a dozen different dimensions, there's no substantially easier method than just one-to-one perfect matching, which requires a very, very large number of patients. You have to take Patient A1, who has subset X1 of 12 different conditions, and find Patient B1 with that exact same subset. Then repeat, 47 times in this case. It is already quite hard to find the person B1 who has the exact subset X1 of conditions that person A1 had. For example, if there are 12 conditions, each condition is present in half the people that come into your clinic, and the conditions are independent, you'll need to go through 2^12 = 4096 people, on average, to find another exact match. The conditions may be a bit correlated but this can only help you so much when you're talking about 12 different conditions. To repeat that feat 47 times is very hard. You'd have to churn through 10s or 100s of thousands of sepsis patients to get your matching subset. This would require access to an enormous pool of sepsis patients and constant reporting of all the conditions they have, that you want to match on. For this to be done without any mention in the paper is utterly beyond belief. And the fact that the counts match near-perfectly in an off-by-1 fashion does not help the situation at all.
- bsdetector 5y agoSee this is fascinating to me. You are the statistician I originally replied to, and yet you also won't clearly answer the questions! This is like pulling teeth, but the implication seems to be my take on this statistical test was basically correct. > Difficulty of matching subgroup counts grows rapidly with number of dimensions. ... and the conditions are independent, you'll need to go through 2^12 = 4096 people [for 12 conditions] Except many of these are "primary" diagnoses, which presumably you have one of hence the name, so it's not possible for a person to have both a primary "Pneumonia" diagnosis and a primary "Other" diagnosis. So for each person maybe you're matching 2^3 or 2^4, not 2^12 - in any case, far less than your conclusions are based on. > You'd have to churn through 10s or 100s of thousands of sepsis patients to get your matching subset. So if the matching was 1/1000th as hard as you thought it was, that would be 10s or 100s of patients. A million cases a year, several doctors in major metro areas, there's probably 10s to 100s of patients at any given time in their hospital systems. > Consecutive means you don't skip anyone who has the condition you're trying to treat. Sepsis is very serious and common, so they'd have to treat them all basically simultaneously with the experimental treatment. I'm no expert, but I'd expect they'd want to be able to abort the trial if people started dying from it. It also doesn't even make sense as a study, because you want to test like for like as much as possible; maybe the treatment works fantastically on patients with cirrhosis and nobody else. I think a plausible scenario is each day before normal rounds they did a med search for a patient matching the next one from first group, generally found a good match, worked with their doctor to change the treatment and added them to the study. Maybe two in a day, or skipping a day, and a month later they have really good matching data. Lots of cases to choose from, many exclusive variables. What do the statistics say for a charitable interpretation? Pretty good odds, right? So maybe they didn't report their methods accurately, maybe it was assumed from domain knowledge or was a mistake when they cut and pasted from a template. Could be fraud too, but that seems like a huge leap to be certain of.