9 ms·
I’m a software consultant for Pharma companies and have worked on projects that used the same kind of study design as the Pfizer COVID vaccine (Kaplan-Meier).
by invokestatic 5y ago
I’m a software consultant for Pharma companies and have worked on projects that used the same kind of study design as the Pfizer COVID vaccine (Kaplan-Meier).
“Raw data” is submitted to the FDA in the CDISC format. This format contains a lot of pretty sensitive medical information, including which diseases a patient has, their medical history, what drugs they take, etc. This is supposed to be anonymized, but if the public were to be able to get this info, I strongly believe there is enough information to re-identify patients. And it’s not as simple as just removing the sensitive medical data because the primary or secondary analyses may be dependent on them.
- rvnx 5y agoAnd ? Where is the bad thing to open a dataset of comorbidities and related / suspected effect of vaccines / medications ?
- jjulius 5y ago>And ? What's not to get? OP said: >... I strongly believe there is enough information to re-identify patients. Their concern is that just dumping raw data like this into the public would be a massive violation of privacy for countless individuals.
- rvnx 5y agoRe-identify based on what ? There are no databases that you can link against if you strip out the major identifiers (like city and date of birth). It's for the benefit of patients, and this raw data is already available, but under a loose NDA for commercial partners and researchers...
- throwawaygh 5y ago> Re-identify based on what ? The fact that you poo-poo this question is telling. Preventing re-identification attacks is incredibly subtle and sometimes impossible. Removing "city and date of birth" is nowhere near sufficient. > this raw data is already available, but under a loose NDA for commercial partners and researchers... Yes. It's available to people who have their real identities tied to their access and can be sued into oblivion or possibly even prosecuted for wanton misuse of the data. Surely you see how this is different from throwing it on a public s3 bucket, right?
- rvnx 5y agoIt looks like from what you are saying that it's an incommensurable effort to anonymize reasonably the dataset (the majority of the risks can be hedge by bucketing, and then there can be subtle deanonymization edge cases but they don't scale). It reminds me the FDA saying they can't release the COVID-19 documents because reviewing 44'000 documents would take 50 years. It didn't take 50 years to produce them... Yes there is some effort needed to anonymize reasonably the data but it's not an impossible task. It's a question of motivation. Here, clearly, the labs and administration don't really want to put efforts into that.
- throwawaygh 5y agoPresumable FDAs analysis assumed current staffing levels. Going to go out on a limb and assume they’ve got a LOT more people involved in creating docs than in deidentifying docs. And, yes, I think public data on this sort of stuff is important and should be properly resourced. I think: Step 1. ASK PERMISSION before changing the way data is handled. Step 2. Share full datasets more generously, but still gated by use and handling agreements. This probably means a private citizens without ethics research training and a supporting it dept can’t get a copy and peruse it on their personal laptop, but also that a truly enormous number of researchers would be able to access the data. The gold standard of a CSV in a public s3 bucket shouldn’t be the enemy of good enough.
- jjulius 5y agoI'm not so sure that you should be so sure that there are no databases that such data can be referenced against. Regarding stripping out major identifiers, please note that OP commented that "it's not as simple as just removing the sensitive medical data because the primary or secondary analyses may be dependent on them". And of course raw data like this could benefit patients, but keeping access to raw data limited also benefits patients. Sure, the data is readily available under a loose NDA, but it's available to people who have explicit knowledge of the requirements around handling sensitive, identifiable data. Average Joe's and Jane's do not have this knowledge, and bad actors just plain don't give a shit.
- ashtonkem 5y agoI mean, it’s not like anonymous datasets haven’t been de-anonymized before[0]. Successfully anonymizing data sets is a hard and subtle process, and it’s way harder than just removing major identifiers like city and birth date. For example age is obviously super important for vaccine and disease studies, and providing an age narrows down someone’s birthdate significantly. And no, this isn’t under a “loose” NDA. This would be covered by HIPAA, which tends to be un-subtle about violations. 0 - https://www.cs.princeton.edu/~arvindn/publications/de-anonymization-retrospective.pdf https://www.cs.princeton.edu/~arvindn/publications/de-anonym...
- hcknwscommenter 5y ago"There are no databases that you can link against if you strip out the major identifiers" Isn't that provably untrue? https://bits.blogs.nytimes.com/2015/01/29/with-a-few-bits-of-data-researchers-identify-anonymous-people/ https://bits.blogs.nytimes.com/2015/01/29/with-a-few-bits-of...
- throwawaygh 5y agoWhat is your full name and address? Age? Weight? Height? Eye color? Blood pressure? What STDs do you have, if any? How often and how much do you drink? Smoke? Any history of drug use? Mental disorders? What about your family? Where is the bad thing in you posting this information here right now? Actually, that's not even analogous. A better analogy would be you post all this info under a pseudonym -- e.g., "rvnx" or something like that -- and then the mods doxing you without your consent.
- cogman10 5y agoI'm 36, I weigh 190 lbs, I'm 6'1. My eyes are green. Blood pressure 120/80 (last I checked). No STDs. I drink once a week about 3 servings. I don't smoke. I don't use drugs or have a history of mental disorders. Depression runs in my family. I only really hesitate to post full name and address but, frankly, I'm sure given my user handle it wouldn't be terribly hard for someone to find both of those. So tell me, what harm have I done to myself by sharing this information?
- jjulius 5y ago>So tell me, what harm have I done to myself by sharing this information? You're comfortable sharing it, that's cool! Is your neighbor? Your friend? Your cousin? Should they be obligated to divulge their own medical information just because you're fine divulging yours?
- throwawaygh 5y agoMore to the point, he didn't even share it. The concern here is re-identification and he chickened out of sharing his identity.
- NicoJuicy 5y agoNo, he gave the path to reidentify him
- 5y ago
- notwhereyouare 5y agodidn't netflix have to stop their algorithm recommendation competition because people were starting to identify real people through the samples provided?
- jacquesm 5y agoSee also: the AOL search CD ROM set and many other releases where it turned out that anonymous data isn't.
- soupfordummies 5y agoI’m intrigued. Got any more info on this or something I can google?
- ceras 5y agoI think the AOL search leak reference is this: https://en.m.wikipedia.org/wiki/AOL_search_data_leak https://en.m.wikipedia.org/wiki/AOL_search_data_leak Even if you take more care than AOL did in anonymizing your data, the unfortunate reality is that any publication of data increases the knowledge an adversary has at identifying somebody. Anonymizing is more about reducing the chance someone is identified than guaranteeing they never will be. And high dimensional data is particularly hard to do so in a way that retains the data's usefulness. 33 bits is an old defunct blog on this topic, but it has some interesting posts and academic papers if you want to go down the rabbit hole: https://33bits.wordpress.com/about/ https://33bits.wordpress.com/about/ Specific paper on Netflix deanonymization: https://33bits.wordpress.com/about/netflix-paper-home-page/ https://33bits.wordpress.com/about/netflix-paper-home-page/
- jacquesm 5y agoThey key is to combine more than one anonymized dataset, this vastly increases your chances at de-anonymization. This paper is a very good starting point: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1450006 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1450006
- nradov 5y agoWhat is the specific basis for your belief that patient re-identification is possible? The federal government has provided clear guidance on de-identification which should eliminate that as a serious concern. https://www.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/index.html https://www.hhs.gov/hipaa/for-professionals/privacy/special-...
- throwawaygh 5y agoNotice that this is a policy document outlining a process, and that one outcome of following that process can be not sharing the data! It's not as simple as "follow these rules and you can share the data". It can also be "follow these rules, which tell you not to share the data". It's also possible to reach an outcome that the data can be shared but must be de-identified in a way that precludes important statistical tests. Any sparse feature is incredibly powerful for re-id, and combined with other features might be difficult to share without running afoul of de-id best practice. The problem: rare medical conditions that must be accounted for in a statistical study of vaccine side-effects are examples of sparse features. So you can share the dataset, but not in a way that's useful for a non-GIGO statistical study.
- rvnx 5y agoThe thing is; in practice, how can you re-id someone that has a very rare medical condition if you don't already know that the person has this very rare medical condition ?
- throwawaygh 5y ago"this person is certainly my biology teacher, who shared the fact that she had a rare genetic disorder during our Genetics unit as well as <insert otherwise non-identifiable columns>. She has herpes. Or also she has an alcohol problem and used to do acid." Teacher is now fired. You also keep ignoring the issue of consent. Step one is to ask patients. Or the person shared their story on the public website of a "Run For The Cure" style website about that genetic disorder. Or so on. "Sparse Features" aren't always a thing that the person wants to keep private.
- mchusma 5y agoI think the level of effort we make as a society on medical privacy is backwards. All health information should be public except in extreme circumstances. Right now, it's almost impossible to get datasets on health data. Imagine if 100% of everyone's data was digitized. The amount of health innovation, lives saved, cures found, is likely to be incredibly large. All the costs that go into trying to comply with HIPPA, could be gone. I get that some people want to prevent the use of the data for discrimination...just legislate that then.
- enkid 5y agoYou don't see any potential misuse if having people's STD results publicly available? (To use an extreme case)
- fxtentacle 5y agoWith black box AI, any legislation on data use is impossible to enforce. So limiting access to the data might be the only way.
- tpoacher 5y agoI hear the concern, but no. There are ways to inject randomness into a dataset, giving subjects plausible deniability without compromising the integrity of the outputs. See https://pair.withgoogle.com/explorables/anonymization/ https://pair.withgoogle.com/explorables/anonymization/ for a nice example. Bottom line is, there are ways to share such data safely. Whether the data owner cares enough to do the extra work, especially if doing so removes a competitive advantage, is another story.
- invokestatic 5y agoI have no qualms with releasing a subset of the data that is properly and irreversibly anonymized. But you have to keep in mind that by introducing “randomness” into the dataset, you limit it’s usefulness to check to make sure the statistics match up with the study’s official topline results. Furthermore, the FDA submission dataset is essentially a database with dozens of tables each with often hundreds of columns. It’s a LOT of data, with exponential complexity to make sure all the right fields are redacted. There’s also the point that pharma companies are under no obligation to release this data. It’s generally considered proprietary. That said, due to the substantial amount of government funding provided to the development of the vaccine, I think we should be entitled to this information.
- ashtonkem 5y agoIf the point is to improve public trust “don’t worry, we manipulated the data to make it more anonymous, but we pinky swear that we didn’t manipulate it in a way that changes the results” is not a good approach.
- tpoacher 5y agoIt is when it's mathematically robust by design, as in the linked post. Nobody's saying "tamper with the evidence to demonstrate favourable outcomes" here. And if tampering with the data to favour the data donor were a concern in the first place, then it should be a concern regardless of whether the data donor said they followed further anonymization protocols or not.
- stakkur 5y agoPHI is not the ‘raw data’ they’re referring to, and you know it.
- ekianjo 5y ago> Raw data” is submitted to the FDA in the CDISC format. This format contains a lot of pretty sensitive medical information, including which diseases a patient has, their medical history, what drugs they take, etc. This is supposed to be anonymized, but if the public were to be able to get this info, I strongly believe there is enough information to re-identify patients At the time when you inject hundreds of millions of people with it, I'd say this should be mandatory to disclose as much information as possible regarding the actual trials. I don't trust the FDA one bit.