3 ms·
The best dataset for this audience is probably the most recent curated set of variant calls on 1000genomes.org. (Incidentally, the data access links appear down
by bthomas 15y ago
The best dataset for this audience is probably the most recent curated set of variant calls on 1000genomes.org. (Incidentally, the data access links appear down at the moment...)
This provides data in VCF format [1], which I would argue is the lowest level you want to go with this data unless you are doing variant calling methods development.
One tool you can use to analyze these data is PLINK/SEQ [2] (disclaimer: I work on the project). If any C++ devs are interested in contributing, let me know. (We'll have the source on Github soon...still pushing my group to learn Git :)
[1] http://www.1000genomes.org/node/101 http://www.1000genomes.org/node/101
[2] atgu.mgh.harvard.edu/plinkseq/
- apaprocki 15y agoSince you are involved with PLINK/SEQ, you might be able to answer this.. Can you recommend any reading for learning how to perform admixture analysis/visualization? I'm interested in seeing the workflow behind starting with raw genomes and producing visualizations similar to Dodecad.
- stopcodon 15y agoI was just having a conversation with my supervisor yesterday about how amazing the PLINK project is, and how much easier it makes my work. I use it every day and I can't count the number of times it's saved me from writing embarrassingly messy R scripts to accomplish something.