11 ms·
Hospitals are selling troves of medical data
- JohnWhigham 5y agoSoon enough this data will end up in the hands of our insurers, and enough technology will be built to where if you buy a 6 pack on the weekend, your premium will be adjusted on-the-fly. It's all so tiring.
- throwaway3699 5y agoFWIW socialised healthcare (the better alternative) hasn't solved that particular problem. Punitive taxes on alcohol and smoking are very common because they're trying to reduce costs. Insurance is the same thing but on an individual basis, which means healthier people should _in theory_ get better premiums in your example.
- JohnWhigham 5y agoRight, but I'm talking about the novel ways companies try to optimize premiums. Car insurance is already doing it with some companies wanting you to install an application on your phone to track your driving habits in hopes of maybe lowering your premium. Disgusting shit.
- pc86 5y agoI don't install these apps but what exactly is wrong with them and why is it "disgusting shit?" If I drive the speed limit or slower, stop for 3-4 seconds at every stop sign, don't accelerate quickly, etc., what's wrong with me having a lower auto insurance premium than someone with an identical profile who goes 10 over the limit everywhere and rolls through every 5th stop sign? I'm not implying insurance companies are doing this out of altruism. If it resulted in a net decrease in profit, the apps wouldn't exist. But it does seem like a beneficial form of price discrimination. It seems like it's only "disgusting shit" if it makes your insurance premiums go up because you're a less safe driver.
- throwaway3699 5y agoI will admit I'm not a fan of surveillance based insurance, either. Just pointing out the alternative is individualising the cost via insurance or making all society pay for a few people.
- kelnos 5y agoThe problem is when it becomes required to submit to this sort of surveillance, or you don't get insurance at all. That doesn't seem like an unlikely endgame to me. Also the factors the app monitors are not necessarily correlated with risk of a crash in the way we'd expect. For example, if you drive the speed limit or slower on most highways in the US, you will be driving slower than the speed of most traffic, and it's more likely that someone will rear-end you.
- Spooky23 5y agoAlready done - this type of data is already used to assess your risk for opioid addiction in several states, especially if you are in a high risk/high cost group.
- unishark 5y agoSo medical science could figure out the health consequence of all your dietary and recreational decisions with amazing granularity, and your take is this is a negative thing because of the possibility of higher insurance rates?
- JohnWhigham 5y agoYes? Are you not aware of how money-hungry these insurance companies are, and how you have to fight them tooth-and-nail to get some things covered?
- missedthecue 5y agoIf you live an unhealthy lifestyle, it's easy to imagine how a fairer deal might cost you more money.
- kelnos 5y agoI'm fine with medical science having this data -- assuming there is a way to ensure they're using it ethically -- but insurers? Hell no.
- unishark 5y agoNo one here is proposing we give this data to insurers. The poster I responded to is proposing we deny medical science this data, and that we remain remain sicker and more ignorant about health, because otherwise the discoveries medical science makes might get used by insurers who somehow gain access to our purchasing data later (another thing no one here is proposing they be given).
- DennisP 5y agoUnder current US law, health insurance companies can't even adjust your premiums for serious preexisting conditions, much less because you bought some beer. Auto insurance could maybe do it though.
- mbg721 5y agoMany plans are highly punitive towards smokers; it's not a huge leap to imagine something new becoming the Sin Of The Week.
- wolverine876 5y agoIt's hard to compare smoking to the 'sin of the week'. It's the leading cause of preventable death in the U.S. (or was a few years ago), and probably has been for generations. It is one of the most studied and prolonged public health issues. It's hard to blame smoking insurance costs on 'sin' - it kills people and increases costs for the insurance company, possibly more than any other choice people can make. If you drive in drag races and ask for car insurance, don't be surprised if it costs more.
- mbg721 5y agoAnd this kind of argument will be made for booze, sugar, meat, etc., but only when piling on the users of each one is politically fashionable.
- JohnWhigham 5y agoDon't think that could not change in a heartbeat, given how powerful the healthcare lobby is in Congress.
- deleted 5y ago[deleted]
- DennisP 5y agoAnd yet the ACA passed in the first place.
- teacup21 5y ago>> Soon enough this data will end up in the hands of our insurers, and enough technology will be built to where if you buy a 6 pack on the weekend, your premium will be adjusted on-the-fly. After watching the movie "I Care A Lot" (https://www.imdb.com/title/tt9893250/ https://www.imdb.com/title/tt9893250/) and reading up the horror stories about forced Guardianship scams in the US (https://www.newyorker.com/magazine/2017/10/09/how-the-elderly-lose-their-rights https://www.newyorker.com/magazine/2017/10/09/how-the-elderl...) -- i'd be worried that corporatized Guardianship companies scan medical data en-masse to find victims (with sufficient work and profit motive, it can be de-anon'd)
- GuB-42 5y agoThe underlying claim here is that de-identification doesn't work, the articles then explores the consequences of that. But the real question should be: why doesn't de-identification work and how to make it work. It is a technical problem, and I thought it was more or less solved. But the author here things it is just placebo. If so, what is the problem exactly? Is there a fundamental problem with the very idea of de-identification? Is there a "bug" in the process? What level is required to conduct these attacks? Can in individual do it? A cybercrime gang? A nation state? Is it only theoretical? Depending on that, the answer could be very different.
- cmiles74 5y agoIt appears there is a spectrum of de-identification tools, some being better than others. https://en.m.wikipedia.org/wiki/Data_re-identification https://en.m.wikipedia.org/wiki/Data_re-identification I suspect that the more data they have for a particular patient (many visits over multiple years) makes the re-identification process easier. The article mention financial data, if dates and amounts of charges aren't masked or altered that could be cross-referenced with another data source to deduce the person's identity.
- prepend 5y agoDe-identification is in the eye of the data steward and their legal folks. So it’s the definition of de-id that doesn’t actually “work” in that it’s possible to reidentify small amounts from de-id data, and that adds up over time. For example, HIPAA considers data de-id if you remove 19 fields or expert determine that it’s de-id. [0] What’s expert determination, who’s an export? That’s up to me to decide and my lawyers to accept. The bug is that if half a percent of people each de-id datasets can be reidentified, likely acceptable in HIPAA, then each data released adds up for reidentified people. And more datasets allow for triangulation and linking to reidentify. The article calls this out as a risk, not a certainty as it’s unclear if anyone is doing this. But the process would be something like: 1) buy HIPAA de-identified data since it doesn’t require patient consent 2) reidentify patients using other data publicly for sale (marketing data, voter registration, etc) 3) new data is not longer HIPAA restricted and fully identified health records can be sold for whatever you like (eg, super targeted drug marketing) [0] https://www.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/index.html https://www.hhs.gov/hipaa/for-professionals/privacy/special-...
- deleted 5y ago[deleted]
- Frost1x 5y agoInformation is information and the more you have, the easier it is to figure out where it came from. Some data makes it easier than others. I worked with a hospital awhile back with patient radiological data (for free and for science). Patients had to explicly sign-off that they were sharing their data with us and what we planned to use it for. A lot of the metadata from DICOM wasn't even wiped, I had their names, street addresses, all sorts of stuff which was supposed to be wiped (de-identified). Even then, I worked with data from a patient with a brain tumor and their neurosurgeon looking to remove it. That data was correctly anonymoized but it was their head, so I could basically reconstruct their face--so is a coarse geomerric sampling of a face de-identified? I guess it depends on how coarse it was. Just look at what's done with social media, advertising, and browser data to get an idea of where things can go.
- deaps 5y agoUnder what circumstances did they "explicitly sign-off" on the data sharing, I wonder? There are a lot of times during a hospital visit, when one could be less-than-observant of exactly what he/she is signing.
- pc86 5y agoEvery doctor's visit I've had in recent memory had a data sharing agreement I had to sign (or at least, that was presented to me) if I hadn't been there before.
- Frost1x 5y agoVery explicitly. This was for a project paternship with a university and hospital to improve patient outcomes with some new exploratory tech approaches. A short document that followed some standard study participation format was generated using easily understandable language in about 1-2 pages IIRC (large fonts so it was easily readable by patients with poor eye-sight). Everything done, including the document, went through an external IRB process for human subject data and was approved. Everyone involved had to go through human subject training and what not. Physician would mention the study to patients that would likely be good subjects for the work about the work, its goals, if they'd be interested in participating. Forms were then provided to patients involved to sign (explicitly) about their agreement to participate in the effort and how their data would be used, protected, etc. The process also required physician sign-off to confirm they read the document to the patient verbally, determined they were competent, cognizant, not under any sort of duress/intoxication, etc. The patient also needed to verbally acknowledge they agreed. Oh, and there was a clause they could retroactively pull out of the work, including their data at any point of they felt uncomfortable or changed their minds. The patients and their data weren't the product, tech developed that would assist patients was the product of the data. For patients who agreed, some would also be permitted to see some of the products of the work related to their data. Im forgetting a lot of the data collection process because it was very rigorous and several years ago now, but everything above bar, no dark patterny ah-ha-gotcha! line buried in a 300 page liability sign off they had to agree to for some necessary life saving treatment or anything of that nature. I even got to meet some of the people we helped which was a bit rewarding to see people's lives improve a bit with technology. The specific patient mentioned and their neurosurgeon even let me sit in on their brain surgery tumor removal (patient's suggestion), which was a very unique experience. So yea, they knew what was going on. With that said, not all data usage was as transparent and ethical as what I worked with, and I saw a lot of mistakes there that make me cringe thinking what a less ethical business with no transparency might do, given the opportunity.
- prepend 5y agoI’m glad to see this getting more attention as it seems scary to me as a private citizen who wants my medical data to stay private. I’m not sure how to defend against this as it seems based on HIPAA and what it allows. Since de-identified data can be legally sold, I think it will. The theoretical defense I’ve thought up is a class action lawsuit for synthetic breaches. Since these data are deemed de-identified by expert determination [0] and that’s hazy, if I could reidentify myself after de-id, and I didn’t authorize it, then I could be eligible for breach damages for HIPAA violations up to $50k per person [1]. Since these sets have millions of people, likely everyone in the country. And since expert determination can possibly classify an acceptable re-id risk as less than 1%, this could be a million or two people. So a big enough pool to attract big legal investments. That would increase the cost and risk of doing this to outweigh the benefits. But currently it’s “free money” for any healthcare system that’s kind of impossible for me as a patient to opt out. [0] https://www.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/index.html https://www.hhs.gov/hipaa/for-professionals/privacy/special-... [1] https://www.injuryclaimcoach.com/hipaa-violations.html https://www.injuryclaimcoach.com/hipaa-violations.html
- aabaker99 5y agoI like your idea of a class action lawsuit but wouldn't the $50k per person penalty shift your idea of what an organization would consider an acceptable risk? 1 million * $50,000 is probably not acceptable. I hope for all our sakes you are in the business of buying de-identified medical data :)
- prepend 5y agoI think this has the potential for the next asbestos or tobacco or opioid payouts. I definitely wouldn’t want to work in this area (both for ethical and business model at risk for sued out of existence).
- notafraudster 5y agoHIPAA has no private cause of action (you can't sue providers who violate your rights under HIPAA). The government can fine them, so from a provider POV they are liable for the breach you propose, but you are not eligible for recovering the $50k. You may or may not have a private civil tort against the medical provider, separately from HIPAA.
- pulse7 5y agoIn a few months somebody can come up with a deep-neural-network to identify those people behind "de-identified data"...
- TrailMixRaisin 5y agoThere is a (famous) quote from Cynthia Dwork, who is most likely the key researcher behind differential privacy: "De-identified data isn’t". You can either re-identify the people behind the data or you alter it that strong that it becomes useless for meaningful applications.
- nradov 5y agoDe-identified data can still be quite useful for some types of research. If you strip away every field of personally identifiable information except, let's say, sex and birth year, there's no way to do meaningful re-identification.
- rscho 5y agoThis is wrong. Rare diagnoses are the simplest case of reidentification, but there are many many other opportunities. There's a whole field of research about that.
- specialist 5y agoDitto time and sequence (order). For example, online movie reviews were deanon simply by correlating viewing history with order reviews were posted.
- tmearnest 5y agoDates and times are generally deidentified by choosing a random initial date and changing subsequent timestamps to the random initial plus the duration between visits. The sequence and delays could potentially be used to identify patients, but this would be a lot harder than having absolute timestamps.
- specialist 5y agoTotally. I recently had a crazy notion for losslessly scrambling the sequence as well. Mostly for protecting voter privacy (order in which ballots are cast). One of the major blockers to fully digital voting. I haven't found any hits using terms like "cryptographic timestamps." Surely I can't be the first.
- aabaker99 5y agoIt's surprising that many medical providers don't understand this side of HIPAA. I recently spoke to a department head of oncology who seemed to think patient consent was required for sharing data and was not comfortable sharing their patients' data. What they don't realize is that HIPAA doesn't require consent if the data is de-identified and so their organization can or is sharing their patients' data anyway.
- SkyPuncher 5y agoThe problem is risk. HIPAA allows for a lot of things that feel like exceptions to the core principles of HIPAA. Many things are vaguely defined as "reasonable" - which changes over time. MD5 was a reasonable password hash - until it wasn't. SHA1 was - until it wasn't. Etc. An article like this can arguably prove that there is no reasonable means of de-identification since multiple data sources can be combined. Combine that with the fact that HIPAA puts pretty high limits on the minimum cohort size that can be associated with a unique identifier and low ROI from actually sharing this data. Many places end up in a position where it's simply not worth the risk to share.
- pitaj 5y agoArguably SHA1 is still reasonable as a password hash, but of course not recommended.
- aabaker99 5y agoOne troubling re-identification attack for medical data is the trail re-identification method [0]. A lot of privacy analysis will consider the data to be in the shape of a table T with some columns A,B,C and use the notation T<A,B,C> to describe a de-identified dataset. The trail method will take multiple de-identified datasets, each from a different hospital, T_1<A,B,C> T_2<B,C,D>, T_3<C,D,E> and use their shared columns to narrow down on a set of individuals. So, even though each hospital may have a legitimately de-identified dataset in isolation, it is not de-identified when combined with the (also de-identified) data from another hospital. The risk of this attack increases as patients visit more hospitals. We humans are fairly long-lived and tend to move around so it may be substantial. (That being said some hospital systems are quite large like Kaiser Permanente and serve huge areas so visiting multiple hospitals doesn't necessarily create multiple tables.) [0] https://dataprivacylab.org/dataprivacy/projects/trails/trails2.html https://dataprivacylab.org/dataprivacy/projects/trails/trail...
- alistairSH 5y agoFor this attack to work, wouldn't one of the tables need to contain PII of some sort? If A,B,C,D,E are all de-identified, the aggregate is still de-identified? But, if E is SSN (or some other PII data), then the entire set can be re-identified?
- taejo 5y agoThat's one option: you combine protected, de-identified information with unprotected (e.g. non-health) information to re-identify the protected information. But also, something like Facebook allows you to target a person who lives in $TOWN, works at $COMPANY, born in $YEAR, even if you don't know that person's name or SSN.
- rscho 5y agoThe funniest part is that this data is actually mostly (and often completely) useless for the stated purposes of statistical analysis. Routine clinical data collection is of abysmal quality, but many buyers don't see the full extent of the catastrophe. It'll be much more useful for legal and not-so-legal purposes. And last but not least: insurance. As an aside, privacy enforcement for medical data is much easier through fingerprinting tethered to a NDA.
- queuebert 5y agoI completely disagree. A lot of the data is standardized already, such as lab values and other so-called "structured" data, and can be used almost immediately without much cleaning. This is most of the routine data collection. The largest remaining part is patient notes, which are free form text but still contain plenty of useful information searchable by keyword or extractable by human curators. You are correct that if doctors could be more strategic about how they collect info it would make the job of the data scientist much easier. This is especially true in radiology and pathology.
- stryker7001 5y agoWhat does more strategic mean? Doctors aren’t data collectors. Their job is to provide patient care. If you want data collected, design a system that doesn't add to documentation burden.
- rscho 5y agoYeah yeah, I've heard the same story many times. So where are the fantastic new discoveries made on mass lab statistics, then? Because epidemiologist have had access to this same data since a long time already, and they are not any less capable than industry "data scientists". The saying goes: "90% of diagnoses are made on patient history". Most of labs require context for interpretation, so the vast majority of what is collected is statistical noise.
- jtaft 5y ago> As long as they de-identify the records — removing information like patient names, locations, and phone numbers — they can give or sell the data to partners for research. I don’t feel this is enough to deidentify. If timestamps or a patient # is associated with records, should be possible to combine with credit card records to discover who someone is. I wonder if we can request what is shared.
- organsnyder 5y agoI work in healthcare (though not with any efforts like those described in the article). Anonymizing data is much more complicated than simply removing the obvious PII: depending on what the data is, even things like procedure codes with dates might be enough to identify a patient. HIPAA and other related regs account for this, and have pretty strict procedures that must be followed before data can be considered officially deidentified.
- stevebmark 5y agoIt would be more honest if the article title said "deidentified" medical data.
- agumonkey 5y agohow come more and more of society is about selling blips of events as data ? is the price of everything far lower than what is required to sustain the system ?
- fidesomnes 5y agocontrolling eye balls is the business model of the 21st century. ads, games, videos, vr, the entire web itself. globalization has brought down the costs of production orders of magnitude lower to produce than what they cost before it. the entire global system is designed on the premise that it is possible to sell to Americans cheap products that requires them to work more and more for new shiny things.
- dillondoyle 5y agoKind of a similarly, I've previously wondered: if we act to shield prescriber data from corps if that would cut down on shady pharma 'bribes' e.g. the opiate crisis and the some other direct to Dr pharma marketing. Insys and Purdue knew which doctors prescribed insane amounts of pills. And then rewarded them with $. Insys even put it on paper as ROI and at least a few went to jail. I'm not sure technically how well (or if legally) it would work though so maybe the answer is not at all. mckesson would still know what pharmacies pills go to and anyone can figure out where a MD works to correlate at least target zip codes/markets. But maybe since we already have this monopolistic distribution setup could prohibit mckesson from disclosing granular shipment data.
- joelbondurant 5y agoHospitals are pure evil.
- yawnxyz 5y agoCould someone here comment on whether it's possible to "remix" and generate faked, synthetic data from real data sets like this, and how that could work? I'm getting my feet wet in the field, but I'm thinking something like the "this face doesn't exist" project but for medical records. I'd love to read a writeup on that!
- fourtrees 5y agoThe author's right about everything but the date this mess started. I worked in Health IT in the early 2010s, and this sort of thing was in full swing by 2013ish. Profitability was clearly the motivation, although the patient (and the true believers employed by the healthcare organizations) were assured that their data would be private, secure, and beneficial to the quality and cost of treatment. I guess the jury's still out on quality.