6 ms·
I've heard about this a lot from people I know who work in healthcare. It seems that one could make a successful business simply by hiring a bunch of data scien
by zo7 10y ago
I've heard about this a lot from people I know who work in healthcare. It seems that one could make a successful business simply by hiring a bunch of data scientists to offer analytics and data processing services for healthcare, but what's preventing that? Is there a lack of expertise, funding, too much regulation, or something else?
- zbjornson 10y agoUnfortunately it's not a simple endeavor. (a) For healthcare data specifically, the data is sensitive. You need to follow HIPAA and usually 21 CFR 11 regulations, and you always face potentially high liability in case of breach. (b) In part for that reason, it becomes expensive. Even for non-healthcare biomedical data (research-only data), most academic labs will not or cannot pay outside firms to do the work, regardless of whether it would be done better or faster.
- lj3 10y ago> Unfortunately it's not a simple endeavor. The HIPAA and other regulations aren't any more annoying than any other modern programming practices these days. As for liability in the case of a breach, that's what E&O insurance is for. It does raise the barrier of entry a little, but it's not by any means prohibitive. I think the bigger issue is what format this data is in. Most medical and medical records data is in a variety of proprietary, non-open and difficult to integrate technologies. One look at HL7 is usually enough to send a programmer back into the loving embrace of anything else. Working with this data is time intensive and expensive and I'm guessing most heathcare companies don't see it as worth the cost.
- op00to 10y ago> Working with this data is time intensive and expensive and I'm guessing most heathcare companies don't see it as worth the cost. It's not worth the cost 99% of the time. The only reason we're still making real breakthroughs is because of the research institutes doing basic research on public funding. Even then, they're swinging for the low hanging fruit.
- lj3 10y agoI read your other comment too. That whole situation breaks my brain. How can you do meaningful long term research if you're not aggregating data? > I couldn't have gotten out of there fast enough for the way saner land of tech companies. You know you have a problem when tech companies seem saner in comparison.
- op00to 10y agoResearchers are aggregating data. There's all kinds of ways you as a layperson can mash up genomic data, and it's easier than ever. The data that's most accessible are well-defined genomic sequences, not cutting edge stuff. Much of what 23 and Me is doing is stems off of these publicly funded datasets. In many cases scientists are required to submit data to journals when they publish. However, data is a funny term! Using genomics data, there are many levels of data that you'll encounter - everything from raw image files (many gene sequencers are actually automated digital cameras, taking pictures of florescent markers attached to the DNA strands) to intermediate sequence files to formatted pieces of selected data. Then, there's the whole toolchain used to go from sample to formatted final analysis? What's required to be submitted? How long should it be kept? Who checks all this to make sure it's not NES roms or Shakespeare instead of the right data? Finally, there's the big questio: how can we be sure that the data captured in the intermediate or final steps of analysis actually originates from the raw data? Should scientists store the raw data (TBs and TB), intermediates (GBs and GBs), or final analysis (MBs). For how long? To answer your question about meaningful long term research - I've personally seen grad students' careers effectively ruined due to shitty data storage hygiene.
- hash-set 10y agoWe're wasting too much $$ on stupid shit like Uber.
- deleted 10y ago[deleted]
- zbjornson 10y ago> regulations aren't any more annoying than any other modern programming practices these days As someone in the field I beg to differ :). And regardless of who pays in the event of a breach, the effects may be sufficient enough to shut down the company for future projects.
- lj3 10y agoCan you expand on that? Everyone I've spoken with about HIPAA say that most companies don't even bother to comply and that nobody is enforcing it. Then again, these people work for companies that are small enough they've never had a breach.
- zbjornson 10y agoThat's rather scary to hear, and I can't imagine that they manage to secure access to any of the major datasets e.g. as contractors for hospitals or insurance companies. You can basically self-certify, but most serious companies will bring in an outside contractor on an ongoing basis to certify compliance. Staff needs to be trained, computers need to be managed, software changes have to be very thoroughly reviewed, updates become slow. It makes it pretty unattractive to enter into for a lot of devs.
- roymurdock 10y agoYou're describing IBM Watson Health. Many other major consultancies have similar analytics offerings for the healthcare sector.
- op00to 10y agoThere's a bunch of issues: - Real, bespoke biomedical analysis is not trivial in effort, cost, or time. There are biomedical analysis systems-in-a-box (look at https://galaxyproject.org https://galaxyproject.org), but that's just canned analysis. To make real breakthroughs, you need rigorous analysis that requires years of experience to be able to perform. - It's easier to get the money to collect the data than it is to effectively steward the data you collect. In a past life, I ran a biomedical research computing facility, and everyone got plenty of money for new sequencers, mass specs, and other fancy instruments. They got plenty of money for collecting all kinds of data. No one would ever add money to their grants to actually STORE the data. They would literally put the data on USB hard drives bought from Best Buy, and left them in file cabinets and on desks. There was absolutely nothing I could do about this, and so I quit. - Research is balkanized to hell. Even though I ran the scientific computing for 20 research labs, each research lab was its own fiefdom. They could decide to obey or disobey my policies at will, since they controlled their own funding. You can imagine what happened when I proposed turning on quotas (~100TB per lab, to start!). Rather than work with my team to determine how to share resources, people would just jump off my high speed facility, buy a shitty cheap JBOD from Dell for their analysis, and store their archives on shitty cheap USB hard drives from Best Buy. The funniest part was that if the hard drive failed, and the data couldn't be restored, in theory the primary investigators could get into real legal trouble. No one seemed to worry. There are a few biomedical research institutes that "get" scientific data stewardship - Broad, Scripps, but for the most part, biomedical research computing is a total clusterfuck and I couldn't have gotten out of there fast enough for the way saner land of tech companies.
- ThomPete 10y agoCould it work as a not-for-profit organization something like Internet Archive or Wikipedia?
- op00to 10y agoIt sort of does already. Not-for-profit organizations like the San Diego Supercomputing Center act as Biomedical-Research-IT-As-A-Service providers, but there's so much competition for grant funding, that if you can get away with doing things cheaper, you will. The scariest part was that before I left, I built a highly scalable, long-term archive for scientific data built on LTO tapes that would allow ridiculously cheap (basically the cost of LTO tapes) on-line and near-line storage. When I left, no one wanted to bother with paying for the upkeep of the hardware, people got bored with swapping tapes, and it eventually died. Oh well. Your tax dollars at work.
- rpedela 10y agoThere are businesses like that already but most (not all) act as consulting services. Beyond regulations, there are three main problems. 1. Each organization has their own data silo and treats it as if it were gold. Therefore it is difficult and expensive to aggregate. 2. Because the data is sensitive, there are often many contractual restrictions beyond HIPAA on how the data can be used. Anything from research purposes only to you can build a commercial product but you can only sell it to the data owner. 3. Even if an organization wanted to share data with reasonable restrictions and pricing, it is often hard to do so because their focus isn't software. So it is difficult or impossible for them to share it. Source: I made a serious attempt at a healthcare data analytics startup. It didn't work out.
- drcode 10y agoThe simple answer is that there aren't many incentives to spend money analyzing the data: On a per-patient basis, the benefits of this sort of analysis are very hit-and-miss and hard to quantify... In the modern world we promote a "one size fits all" health model that is encouraged by the many third parties involved in the doctor patient relationship (third parties such as insurance companies, employers and government regulators). It's going to be very difficult to adopt this model to the patient-tailored healthcare that is going to be required to fully leverage the recent advances in machine learning.
- randycupertino 10y ago> The simple answer is that there aren't many incentives to spend money analyzing the data Except for when you want to cherry pick through it to help get your research paper published.