5 ms·
Large Scale Medical Data Mining research, similar to OpenAI. Specifically Computational Healthcare a Search and Aggregation Engine for Medical Records & Claims.
by visualsearchsv 11y ago
Large Scale Medical Data Mining research, similar to OpenAI.
Specifically Computational Healthcare a Search and Aggregation Engine for Medical Records & Claims.
We believe that this is a classic Software eating the world situation and the time is perfect for it.
Here is the link http://www.computationalhealthcare.com http://www.computationalhealthcare.com
We have access to almost 130 Million de-identified medical records from approximately 36 Million patients (~10% of US population) this includes all Inpatient, ED, Ambulatory Surgery records between 2006-2011 from California. To put simply if you lived in California and went to a hospital there is 95.9% chance that we have your data. This data has been available for quite some time but its use has been hindered due to lack of good software. The data has led to significant research, e.g. my collaborator (not me) published a paper showing risk of strokes following pregnancies in New England Journal of Medicine last year.
At Cornell Tech & Weill Cornell Medical College, we have developed a Search and Aggregation engine that will revolutionize how researchers and physicians use this data. Imagine your mother with Leukemia in Remission just got admitted for Pneumonia. With our software, the Physician will be quickly able to asses likelihood of this occurring and rule out any confounding adverse events. Or consider that there is a rare combination of diagnosis e.g. Graves Disease and Clotting disorder that is indicative of a unique genetic mutation likely to offer novel insight into disease process. With our software questions like these can be answered within second, Today & Right now.
The Data, Legal structure and fully functional prototype are available right now. We were counting on support from AHRQ, but sadly the agency has run into trouble due significant budget cuts.
- johnloeber 11y agoOut of all the proposals I'm reading, this seems to be one of the most realistic, for one reason only: this is something where a little money and a lot of computer science knowledge can go a very long way.
- AlexDanger 11y agoI agree this sounds a great candidate for YC. I hope you are looking to take this global. The cost/benefit in health outcomes would be enormous. Most importantly, I think a system like this must be non-profit and freely available for global usage. Privacy issues can be solved with careful ETL process. Ideally a single global platform...although political realities may force multiple instances. Regardless of implementation details - this seems to fit the YC Research charter in terms of medium/long horizons and big 'change the world' payoffs.
- visualsearchsv 11y agoThanks, Alex. I agree that such system should be global / non-profit & free. We have also studied privacy issues surround such system in detail. A large motivation is that regardless of the privacy preserving technique employed the system is a huge improvement over current practices which involve sending out entire data to each individual research group. Today a physician suspecting novel association e.g. Adverse event particular to a co-morbid condition. Usually has to wait multiple weeks for Ordering data from US government, Developing SAS scripts and finally conducting the study. Wrth our stack the underlying question can be answered within seconds. With enough legal permissions we can modify it to deliver only the required data, for further statistical analysis. Sure such a system might not be completely public, but there is nothing that prevents us from giving access to say all board certified physicians globally.
- AlexDanger 11y agoI agree this sounds a great candidate for YC. I hope you are looking to take this global. The cost/benefit in health outcomes would be enormous. Most importantly, I think a system like this must be non-profit and freely available for global usage. Privacy issues can be solved with careful ETL process. Ideally a single global platform...although political realities may force multiple instances. Regardless of implementation details - this seems to fit the YC Research charter in terms of medium/long horizons and global 'change the world' payoffs.
- zo1 11y ago>"We have access to almost 130 Million de-identified medical records from approximately 36 Million patients (~10% of US population) this includes all Inpatient, ED, Ambulatory Surgery records between 2006-2011 from California" Not to put a spanner in your works, as I agree your goal is worthy of a lot of effort from the general public. But, how exactly did you get access to this data, and is it legal? Additionally, from my point of view, even if my medical "records" were de-personalized and made anonymous, I still want to have full control over who/what get's access to it.
- yummyfajitas 11y agoBy suitably modifying queries to make them "differentially private" (technical term), one can allow queries on the data set which have an arbitrarily low probability of releasing personal information. http://research.microsoft.com/pubs/74339/dwork_tamc.pdf http://research.microsoft.com/pubs/74339/dwork_tamc.pdf Building a database which can be locked down to differentially private primitives (differential privacy composes) would allow researchers to partially unlock this data while ensuring that your medical records are private. As long as we choose epsilon (the privacy parameter) sufficiently low there is then no ethical need to ask a bunch of fickle data points for permission.
- visualsearchsv 11y agoWe do something similar. We precompute/aggregate exhaustively by following certain aggregation strategies. The aggregated statistics are further processed to ensure privacy. Differential Privacy cannot be directly applied since the underlying assumptions are too strong. An important consideration is that the error/noise added is independent of the answer. Which means that the system becomes unusable for almost all queries other than general trends. By restricting the query structure, we no longer need large amount of noise. Privacy of hospitals and providers is also very important and cannot be encoded in the Differential Privacy framework. Again this is still a hotly debated issue. But even the most vocal supporters of differential privacy agree that it might not be directly applicable for healthcare domain. Following are some of the paper that discuss this: http://www.openu.ac.il/personal_sites/tamirtassa/download/conferences/anonymity_dp.pdf http://www.openu.ac.il/personal_sites/tamirtassa/download/co... http://www.jetlaw.org/wp-content/uploads/2014/06/Bambauer_Final.pdf http://www.jetlaw.org/wp-content/uploads/2014/06/Bambauer_Fi...
- Thriptic 11y agoIs this data publicly available or have you made arrangements with a hospital group?
- visualsearchsv 11y agoThis data is available through Agency for Healthcare Research & Quality HCUP project. Getting access to it is straightforward if you are affiliated with a teaching hospital and/or a university. There might even be some researcher at your institute who already has access to it. Getting access to it as a private entity (E.g. a startup) is more challenging and often requires a stricter review.