4 ms·
It is a good idea in general, but the idea of infosec for genetic data is a minefield. I deal with similar limitations daily -- one of my areas is large-scale m
by xaa 11y ago
It is a good idea in general, but the idea of infosec for genetic data is a minefield. I deal with similar limitations daily -- one of my areas is large-scale meta-analysis of expression data, which was dandy when that data was collected using microarrays. Now, it's RNA-seq, so a lot of that data indeed stays locked up and you have to apply for special permission to access individual datasets from dbGaP, making large-scale studies difficult/impossible.
But although "algorithms in/results out" sounds good in principle, I think it will be hard to implement in practice. You would have to make algorithms run without network access to prevent a bulk_send_data_to_ip() type of function from being written, but that would hamper complex programs requiring external data.
In general I think the only realistic way forward is to take the 1000 genomes approach of finding people who are willing to take the privacy risks of truly open-sourcing their data. But it sounds like an interesting idea and I hope I'm wrong and your approach turns out to be workable.
- subcosmos 11y agoI like your thoughts here! Indeed I have been hoping to mostly attract people who are willing to be fully open with their data. If I ever get big enough to implement this 'algorithms in/results out' approach I intend to re-engage the whole userbase and have people opt-in to crowdsourced scientific analysis. To prevent data leaking I figured we would start with full code review, and indeed air-gapped analysis. Its a continually evolving thing. I imagine it would be years before I get to that stage. Depends on if I find funding or university help. Hit me up at info@infino.me if you'd like to chat more.