4 ms·
I have a friend who is a medical researcher and it definitely seems they are stuck in the past. In order to study something, he has to: * Come up with a hypot
by jpobst 9y ago
I have a friend who is a medical researcher and it definitely seems they are stuck in the past.
In order to study something, he has to:
* Come up with a hypothesis that X may cause Y
* Request access to data about that hypothesis
* He is only given the data regarding his hypothesis
* He can then study whether his hypothesis has merit or not
We should be dumping these whole datasets into machine learning and having computers give us potential links to explore. Obviously there will be plenty of things that turn out to be unrelated, but it's also very likely the computer can find links that a human would not have considered.
I don't see it changing any time soon in the US, but I suspect other countries with this data will use it, and we'll find the next generation of medical breakthroughs no longer come from the US.
- blaurenceclark 9y ago100% agree I hope we can get there!
- diegoprzl 9y agoWell, that's the way it's meant to be. Exploratory analysis doesn't have the same purpose as hypothesis testing.
- troyastorino 9y agoSadly, even countries with universal healthcare systems don't have universal health informatics systems (the NHS is a prime example — they spent £12B trying to build an integrated system [1]). Lots of countries attempt, including the US — HIPAA was actually originally about data portability [2], and we just spent another $40B [3]. Thus far only smaller countries have had success with integrated health IT systems [4]. [1] https://en.wikipedia.org/wiki/NHS_Connecting_for_Health https://en.wikipedia.org/wiki/NHS_Connecting_for_Health [2] https://en.wikipedia.org/wiki/Health_Insurance_Portability_and_Accountability_Act#Title_I:_Health_Care_Access.2C_Portability.2C_and_Renewability https://en.wikipedia.org/wiki/Health_Insurance_Portability_a... [3] https://en.wikipedia.org/wiki/Health_Information_Technology_for_Economic_and_Clinical_Health_Act#Meaningful_use https://en.wikipedia.org/wiki/Health_Information_Technology_... [4] https://en.wikipedia.org/wiki/Healthcare_in_Denmark#eHealth https://en.wikipedia.org/wiki/Healthcare_in_Denmark#eHealth
- herman5 9y agoIt's frustrating that this problem is more of a political labyrinth than a technology endeavor.
- troyastorino 9y agoDefinitely true. In the US, because fax and phone were carved out in HIPAA as not being ePHI, they have this special protected status that makes it totally cool for providers to fax records around, but sending emails something that risks jail time :/ I wouldn't underestimate the technological barriers to making interoperable health record systems actually useful. There are a lot of different kinds of medical information (SNOMED CT, the best ontology for healthcare, has >1M concepts!), and the best way to structure that information is an unsolved problem. There are lots of different ways out in the wild (complicated by there being lots of half-assed EHRs that were just made to grab incentive money), and the standards that are out there don't really help things (they are so broad that basically every EHR implements their own "flavor" of the standard).
- kharms 9y ago>We should be dumping these whole datasets into machine learning and having computers give us potential links to explore. You're describing P-value hacking, thus named because hack scientists use this technique to publish papers about nonsense.
- highd 9y agoYou basically just need to reduce the P-value you need to claim significance (0.1 to 0.001 or less) to account for the probability of finding those correlations even in noise. This is part of why particle physics has such high standards - you can find a lot of things in the TBs of data CERN generates. See for example: https://en.wikipedia.org/wiki/Genome-wide_association_study https://en.wikipedia.org/wiki/Genome-wide_association_study There's a figure in there depicting associations with P-values of 1e-8: https://en.wikipedia.org/wiki/Genome-wide_association_study#/media/File:Regional_Association_Plot.png https://en.wikipedia.org/wiki/Genome-wide_association_study#...
- specialist 9y agoEvery usage of every bit of medical data requires patient consent. The potential for abuse is not hypothetical. While I was implementing medical information exchanges, every single participant considered patient data to be their own, to be used as they wish. Our (grand)parent company, a lab, was negotiating with Microsoft, Google, pharmas, etc. Each was trying to figure out how to monetize it. For example, targeted ads. The C (executive) level players mocked HIPAA and the other (meager) patient and consumer protections the same way they mocked Sarbanes-Oxley, environmental protections, financial reporting requirements, etc. If you think Google and Facebook are bad... --- My data, all that is known about me, is my identity. It's me. At the very least, if someone's going to profit from my data, I want my cut.
- mattjack 9y agoI agree with kharms >You're describing P-value hacking Here's an example of what can happen when you take a huge corpus of data and throw an equally huge number of hypotheses at it to see what sticks: https://io9.gizmodo.com/i-fooled-millions-into-thinking-chocolate-helps-weight-1707251800 https://io9.gizmodo.com/i-fooled-millions-into-thinking-choc... tl;dr: he "proved" chocolate causes weight loss by comparing chocolate- and non-chocolate-eaters on a very high number of health indicators. That also introduces the multiple testing problem: https://www.wikiwand.com/en/Multiple_comparisons_problem https://www.wikiwand.com/en/Multiple_comparisons_problem The more statistical tests you run against a set of data (EDIT: the more variables you test against a dataset), the higher the chance you get a statistically significant result from random error alone.
- devrandomguy 9y agoIANAS, but does this mean that a set of raw data loses value, as more information is extracted from it? If I use your old raw data to validate my hypothesis, does that somehow also weaken the statistical evidence for your hypothesis? I really need to go back and study statistics, this is getting embarrassing.
- mattjack 9y agoI worded my comment incorrectly (and edited it accordingly). What I should have said is that when you run a stats test against a dataset, there's a known probability that you'll get a significant correlation simply due to chance. The more variables you examine, the higher that chance becomes. I just found this on Google but the first page of this paper explains it a little better: http://www.stat.berkeley.edu/~mgoldman/Section0402.pdf http://www.stat.berkeley.edu/~mgoldman/Section0402.pdf
- thaumasiotes 9y agoIt means that you can't use the same data to confirm a hypothesis as you used to generate the hypothesis. Defensible statistical practice would be to throw anything you like at the original data set, come up with whatever ridiculous idea, and then collect a new data set for the purpose of investigating your ridiculous idea. The original data set provides zero[1] evidence for a hypothesis that it inspired you to think of. [1] Not really, but this is the cleanest way to sidestep multiple comparisons.