2 ms·
I'm imagining some kind of machine-learning training set based algorithm for determining the ancestry percentages. Even at 99.6% difference (due to mutations ov
by captncraig 8y ago
I'm imagining some kind of machine-learning training set based algorithm for determining the ancestry percentages. Even at 99.6% difference (due to mutations over time or sampling error) over 700k samples there's gonna be ~300 sequencing differences between twins. Enough to make their ML system draw different inferences for sure.
- buboard 8y ago23andme has some publication about their method. the thing is , they have a lot more data than what is available in public repositories so they basically established their own system.