3 ms·
Been talking/planning it for a few months but broke ground over the holiday break. I'm working on creating an accelerated genomics data pipeline. Core focuses
by rfc 11y ago
Been talking/planning it for a few months but broke ground over the holiday break.
I'm working on creating an accelerated genomics data pipeline. Core focuses is on the following:
1) Accelerating the genome alignment process from ~45 min to <1 min. (completed)
2) Database that allows users to compare up to 10 genomes for their differences/similarities with basepair mutations (in progress)
3) Ability to overlay data science-y stuff (H20 framework) against up to 10 genome mutation datasets to run clustering algos (or other methods). (Q2/Q3 2016)
I just got a basic prototype up and running that overlays a visual interface on top of a full human genome mutation data set. Currently wrapping up a disease comparison table so that user can run intersections to find potential areas of diseases.
There's lots of tools out there that do stuff like this already but they're often in terminal, require lots of coding exp. or stats, and are really technical to use. Additionally, lots of challenges on the data front since these data sets can be between 3gb to 150gb per human genome. This makes doing comparative genomics at scale hard.
My hope is to build a basic prototype that is visually easy to use, straight forward, fast, and provides proper insight. The ultimate goal would be to create a platform that allows users to do massive population-based genomics (10,000 genomes+). Reason? Because it's super fascinating and once I get my genome sequence, it will be really fun to dive into the software of me.