14 ms·
Sequencing your DNA with a USB dongle and open source code
- LinuxBender 5y agoThis is very cool. Are there by chance any associated projects that could evolve into something like 23andme but remain entirely within a private network meaning that the data is entirely in the hands of the individual?
- netizen-936824 5y agoSounds like a fediverse project?
- Malp 5y agoOh God, I would not want a distributed group of actors with limited trust to sequence my DNA. Maybe it's a project for close group of friends that would be interested?
- netizen-936824 5y agoI wasn't thinking sequencing but rather comparison. Could even hash data for comparison to enforce privacy (unsure how effective that would be) But this could enable things like finding relatives which is what I got out of the comment about 23andme. Instead of all the data being centralized, storage and comparison could be distributed
- Malp 5y agoAh, thanks for clarifying. I misunderstood the main idea behind your comment.
- inciampati 5y agoYour DNA is almost exactly the same as other people's, just a unique mix. Not sure what you are concerned about. What would you expect a bad actor to do with your DNA sequences? I'm genuinely curious.
- fragmede 5y agoA very practical reason not to want your DNA out there, unrestricted, is insurance costs. From car insurance, to health insurance, to mortgage lending rates, and life insurance, and while GINA from 2008 is supposed to protect that information, there are loopholes with the interpretation of that law that should give everybody pause.
- snovv_crash 5y agoUsing that analogy, all the 1s and 0s in your private key are the same as everyone else's as well. Genetic data can be used for all kinds of things, the worst of which would be things like targeted diseases or planting your DNA at a crime scene.
- inciampati 5y agoActually it's like your private key is made up of ~1000 1 mb pieces that each have 1/1000 rate of difference with any other similar piece. Oh, and the order of the pieces is almost always exactly the same. No, genomes are not "almost the same" because they are all in base-4 sequences and this made up of the same 0s 1s 2s and 3s. We are astoundingly similar, even unusually so for a large mammalian species.
- LinuxBender 5y agoYour DNA is almost exactly the same as other people's, just a unique mix. Music is exactly the same notes, just a unique mix. So why is Sony upset that I want to stream their entire library? But jokes aside... A few decades ago I fought the military on collecting my DNA. I stalled them long enough to get my honorable discharge and avoid that all together. It's funny you ask because the commander asked the same thing and joked "Are you afraid we are going to clone you?!" to which I replied, "No sir, you should be afraid you are going to clone me." and we both had a laugh because he knew I was right. The military are not fond of critical/free thinkers. One of me was plenty. I explained that insurance companies were already using this data to retroactively cancel peoples policies even if they were not actively afflicted by something. The commander showed me how to use the FOIA request system. Laws have evolved a little since then but there are plenty of other risks. For starters, I can't easily change my DNA like I can change my debit card. That data can be used to tie me to others or guilt by association which is undesirable drama. It can also be used to try to sell me things. It can also be used to target biological weapons against specific groups of people. There appears to be an imbalance of data sharing in this regard. [1] Then there is simply the matter of privacy. If I want to share my DNA with some lab that is in turn going to sell it out to hundreds of other companies over and over forever, I should at very least be getting paid a vast amount of money and land and have legally binding contracts and NDA's that cover what is and is not allowed to be done with my data and how long it may be retained. That contract and the laws enforcing the contract must have some serious teeth with very serious ramifications for anyone violating it whether intentionally or by mistake. [1] - https://www.youtube.com/watch?v=biNxl7tiVSY https://www.youtube.com/watch?v=biNxl7tiVSY
- mylons 5y agoyes. if you wanted to annotate your genome you could “easily” do it on your brand new macbook (this is ram intensive, you probably need 32G). you’d need a reference genome, like https://www.nist.gov/programs-projects/genome-bottle https://www.nist.gov/programs-projects/genome-bottle then you’d need a program like bwa http://bio-bwa.sourceforge.net/ http://bio-bwa.sourceforge.net/ to map your data. then use https://samtools.github.io/bcftools/howtos/variant-calling.html https://samtools.github.io/bcftools/howtos/variant-calling.h... or something else to produce variants from the mapping results. then compare your resultant vcf file to something like dbSNP: https://www.ncbi.nlm.nih.gov/snp/ https://www.ncbi.nlm.nih.gov/snp/ at this point you can start generating a raw version of a 23andMe report.
- LinuxBender 5y agoNice! Thankyou for the links. I will research all of this.
- mylons 5y agogood luck! it’s not that tough, just a lot of new vocabulary.
- tootie 5y agoI'm unclear from this what kind of equipment you need to extract and analyze the material?
- mylons 5y agoyou’d likely to have to get the nanopore sequencer in the article or find a lab using Next Generation Sequencing to sequence your DNA and give you “raw data” which are usually fastq files
- pas 5y agoCould you please explain how this mapping works? Why it needs so much RAM? Is it doing a fuzzy search of sorts for known sequences (genes)? Why can't it do so one by one?
- ampdepolymerase 5y agoA used laboratory grade NGS system can be had for less than 10K https://www.ebay.com/itm/265148387179 https://www.ebay.com/itm/265148387179 Nanopore is still not quite ready yet for precise and high accuracy sequencing. Give it another five years.
- mylons 5y agowow i didn’t know they were that “cheap” now. i used to work for a major competitor to the sequencer you linked, the SOLiD. and i feel like nanopore is the VR of dna sequencing. it’s always just another few years off.
- joshuamcginnis 5y agoWhat do you mean by it's always a few years off? Nanopore will allow you to do high-quality genomic sequencing _now_, in a home lab if you wanted, for less than $3K. If you amortize the 3K by the number of genomes you can sequence on the same flow cell, the price per base or per genome falls precipitously, depending on the size of the genome of course.
- ampdepolymerase 5y agoThe one I linked to is a decade out of date and OEM discontinued.
- mylons 5y agoya my first thought was how hard are reagents to get, but probably not that hard. i wasn’t in the lab, i was in bioinformatics so i’m generally clueless on reagent acquisition.
- divbzero 5y ago> and i feel like nanopore is the VR of dna sequencing. it’s always just another few years off. Is this also true for nanopores in protein sequencing? This HN comment from a few weeks back [1] pointed out recent progress but perhaps the tech is still not quite there. [1]: https://news.ycombinator.com/item?id=29481075 https://news.ycombinator.com/item?id=29481075
- kingcharles 5y agoSo, how long before I can take my DNA "ROM" file and boot it in an emulator that would allow it to grow?
- Lev1a 5y agoAn idea just popped into my head reading your comment: What if you could take the (binary) data file of your DNA and use it as input in the (recently remastered) Monster Rancher games to generate a monster? Apparently those games use external user-provided data (like music CDs, game discs etc.) to generate the monsters the player would then train and use (something I only recently learned about through gaming livestreams). I'd actually like to see the level of jank that would come out of something like that.
- dekhn 5y agoit's unlikely we would ever be able to achieve this. Even simulating a single cell at high resolution is a serious challenge.
- 323 5y agoYou seriously underestimate the continuous growth of computer power. And quantum computers after, which are perfect for simulating chemical reactions. What was unthinkable 50 years ago, playing chess better than a human, it's now trivial for a $100 device. And it's not necessarily required that to simulate the growth of a human you'll need to simulate the entirety of chemical reactions in all 50 trillion cells and all that.
- dekhn 5y agoIt's possible I underestimate, but I have worked in all the relevant fields of simulation, ~20 years of running various simulations on large HPC, built the largest instance of folding@home using idle cycles inside google data centers, published papers simulating proteins, developed infrastructure to process the voluminous data, etc, etc. Quantum computing remains fantasy (in terms of being useful for science). It's unlikely even if we improved computing hardware many orders of magnitude beyond all reasonable predictions, that the calculations would be able to simulate all the necessary details; most of our simulations now are based on many approximations due to hardware limitations. As to the question of "what level of fidelity is required to turn a FASTQ of somebody's genome into an accurate model of the resulting human, with some sort of realistic environment also provided", that's so far beyond what is even remotely comprehensible it's not worth speculating about in terms of science fact; it's just fiction.
- m12k 5y agoI'm really curious about what I could learn by getting my DNA sequenced, but I'm worried about my rights to not have it recorded and shared without my consent if I got someone else to do it for me - so any advance toward an affordable home test setup is very welcome.
- biophysboy 5y agoIts only valuable if somebody also interprets it for you, such as telling you whether you have a genetic predisposition for certain diseases.
- DoctorOW 5y agoIs that not something software can theoretically provide?
- jacquesm 5y agoYour DNA can tell you a lot about what could happen, but not about what is happening.
- m12k 5y agoOne of the other comment threads indicates that the data, that you need to do that kind of annotation of the sequence, is to some extent available for home use as well: https://news.ycombinator.com/item?id=29695449 https://news.ycombinator.com/item?id=29695449 I'm really hoping someone will work on an open source "23andme@home" solution that ties all this together in an accessible way.
- rumblerock 5y agoYears ago I used Ancestry, then requested the .txt file and asked them to delete it from their records. Uploaded it to run a report at https://promethease.com/ https://promethease.com/ that cross-references your SNPs against the existing body of genetic research. The results have been pretty astounding. I found markers that pointed to poor response to a specific blood thinner my grandfather was put on before he passed. Currently I'm researching the cluster of Bipolar / ADHD / SAD symptoms I experience that all seem to trace back to a certain genotype of circadian rhythm genes I have (thank you, Sci Hub). To boot, some of the studies I've come across have been done on Han Chinese populations that match my descendance. Perhaps going too far down this rabbit hole poses a self-diagnosis risk, but the correlations to my family history and my own life experience working with doctors to diagnose and treat symptoms are pretty undeniable. And given that your run-of-the-mill psychiatrist is going to treat you off of a DSM checklist, I feel much more confident knowing there have been genomic studies to back things up, since my doctor isn't up to date on this research, and finding one that would be will be difficult and expensive. I've shared the papers with my doc and he's been supportive, sometimes I feel like I should be getting a discount on services rendered.
- fragmede 5y agoI don't know if this is the exact nanopore USB dongle used in the article, but this one is $1,000 for the base package, first released in 2014 https://store.nanoporetech.com/us/minion.html https://store.nanoporetech.com/us/minion.html https://www.extremetech.com/extreme/190409-minion-usb-stick-gene-sequencer-finally-comes-to-market https://www.extremetech.com/extreme/190409-minion-usb-stick-...
- koeng 5y agoYep that’s the one. They update the flow cells over time. The bit they don’t tell you is the stuff you need, like a qubit, to properly run the thing.
- joshuamcginnis 5y agoA qubit or fluorometer isn't required. You can use a simple DNA ladder to measure the relative quantity and quality of DNA that's good enough for nanopore sequencing. I just did a full genome sequence of a novel fungus using this exact approach.
- koeng 5y agoHuh, interesting. Did you fragment? I’d imagine comparison of high weight gDNA wouldn’t be too nice on a gel. You also still, in that case, need a gelbox + ladder + loading dye + sybrsafe or whatever, so it’s still not nothing.
- joshuamcginnis 5y agoI did a HMW extraction kit on the DNA and used a gel to estimate the volume of HMW DNA. Yes, you need to be able to run a gel, but I'm not sure what the expectation is from folks; that you just place a random piece of non-sterile tissue on a chip and have it do the extraction, sequencing and assembly? That seems like an unrealistic expectation.
- 5y ago
- GekkePrutser 5y agoI don't see any reference to the "USB dongle" mentioned in the title. I was thinking this would be some cool thing you could do at home.
- dekhn 5y agohttps://nanoporetech.com/products/minion https://nanoporetech.com/products/minion
- GekkePrutser 5y agoAh thanks! Not something 'just for fun', so. But good to see this tech is becoming more affordable!
- inglor_cz 5y agoDNA sequencing bugs me quite a bit. On one hand, I would love to learn something new about my body. On the other hand, what if the results tell me that I am predisposed to some horrible untreatable disease? Will I spend the rest of my days observing every little pain or discomfort and thinking "is this IT?"
- nomercy400 5y agoHow about affinities to possible health issues, which could be avoided if you started now and not in 20 years?
- inglor_cz 5y agoI know. There is a lot of different scenarios. It is the worst one that bugs me. Human nature in action. Perhaps a trusted middleman would be a solution: "just don't tell me about anything that is totally beyond my control".
- wallacoloo 5y agowell, build a whitelist of the conditions you are interested in knowing. then just run the report through a sed filter so that it strips out all the information you’re not interested in. destroy the original report. problem solved: infohazards avoided.
- monopoledance 5y agoI think you would have to two scenarios at hand: 1. A completely genetically determined disease; a rare 100%-going-to-happen deal. (Which you would probably know about already, because your mother, or grandfather died from it...) 2. Some significant, but abstract risk modification. With 1., you would know, you will get sick/die some time soon in the future, allowing you to live your life accordingly, die without regrets, prepared and so on. You can take that into consideration when planning for a family, taking job offers, procrastinating on the good life with work and retirement plans. Burn bright. With 2., there is a very, very high chance lifestyle choice influence the stated risk, as obviously not everybody who got the polymorphism gets sick. So you can get your ass up, exercise, quit smoking and drinking, reduce stress, get regular check ups, ..., and avoid getting sick or reduce the impact/progression, in case you do. I think, logically, knowing is always better than not knowing. But I understand how anxiety does tell a different story.
- dekhn 5y agoFolks are free to analyze my genome, https://my.pgp-hms.org/profile/hu80855C https://my.pgp-hms.org/profile/hu80855C Last time it was analyzed the conclusion was that there was nothing actionable.
- zmmmmm 5y agoHave you ever encountered any insurance implications from it? eg: questioned whether you have ever had a genomic test etc. and had to answer yes and then them wanting to see results? I guess in your case where nothing actionable is found it's benign. It will be the cases where there are risk factors for late onset things - cancer, diabetes, heart disease etc. where it would get sticky.
- dekhn 5y agoNo, my health insurance company doesn't care about my whole genome data. Health Insurance companies are already quite skilled at (and profitable due to) their ability to model life expectancy and health issues without genomic data, and they are legally prohibited from using this data, in my country anyway. Life insurance is different (they are allowed to incorporate much more information) but I've never been asked for anything like that. As for the case where nothing actionable is found- it's not benign. It's absence of information, not information of absence.
- Cyclical 5y agoNanopore sequencing is a really interesting technology. It utilizes fundamentally the same apparatus as a Coulter Counter [1], which is a general method of counting and sizing arbitrary particles that's frequently used in flow cytometry. Applying it to sequencing by drawing unwound DNA through the pore was a really excellent logical leap, and we're only now starting to see the benefits of even though it was first ideated over 30 years ago. [1] https://en.wikipedia.org/wiki/Coulter_counter https://en.wikipedia.org/wiki/Coulter_counter
- billiam 5y agoTMI.
- a-dub 5y agothe nanopore units are awesome! although if i recall, most of the device is a replaceable one time use consumable and the cost of that consumable is quite expensive (at least hundreds, if not thousands). when i looked i was interested, but was turned off when i saw that the cost far outstripped commercial sequencing services.
- lend000 5y agoHow does it get the DNA to go through the hole?
- Cyclical 5y agoInitially, the DNA is brought near the pore through diffusive (brownian) motion + any small attraction it'll have to the membrane. Close to the pore it uses a combination of the electrophoretic and electro-osmotic effects to draw the DNA molecules through. The application of an external magnetic field will cause the charged DNA molecules to migrate along the field (electrophoresis). This is independent of the fluid, and happens to any ions under voltage. The electro-osmotic flow, on the other hand, is a motion of the fluid itself, pulling the DNA molecules along with it. EOF is a really interesting phenomenon which is caused by the interaction between the surface chemistry (vis-a-vis charge distribution) and the concentration gradient of charge carriers in the fluid. I'd recommend Fundamentals and Application of Microfluidics by Nguyen et al if you're looking for a good primer on electrically induced flows in microfluidics.
- thadk 5y agoMaybe at our local library we should be able to check these nanopore sequencers, or even other devices like simple & robust medical devices like handheld ultrasound devices that plug into iPad's?
- wombatmobile 5y ago> Why not make the software into a proprietary product? ... There’s such a race there that it’s hard to commercialize the software for the long term.” Schatz continues, “Plus our work is largely funded through government sponsored grants, so this is one of the important ways for us to give back to society.” In some people's thoughts, making a better society is the first and most obvious thing to do with technology like this, not an accidental consequence of inconvenience. Fortunately, enough of those people are active in the world to make Main Street different to Wall Street, at least sometimes.
- klmr 5y agoIt’s a weird quote anyway since there is commercial, proprietary software for DNA sequence analysis. Just a few examples of companies in this space are Sentieon, Edico (acquired by Illumina) and Parabricks (acquired by Nvidia). And Michael knows this (they’re sufficiently well known, and his own research laid some of the earliest foundations that Parabricks would ultimately build upon) so I’m assuming the quote was taken out of context or he was talking specifically about his own lab.
- up6w6 5y agoReports of people trying to use it at home without any special lab: https://abarry.org/dna-sequencing-in-our-extra-bedroom/ https://abarry.org/dna-sequencing-in-our-extra-bedroom/ http://blog.booleanbiotech.com/sequencing-at-home-with-flongle.html http://blog.booleanbiotech.com/sequencing-at-home-with-flong...
- twotwotwo 5y agoA researcher mentions using a compact index based on the Burrows-Wheeler Transform to fit things in less memory compared to using a huge hashtable. I see open-source implementations of BWT-based indexes (FM-Index/FMtree) out there. Out of curiosity, does anyone know of anything using BWTs for compact indexes in more everyday uses (like full-text search), or alternately reasons it doesn't really work outside the genome-alignment use case? Likely it only 'pays for itself' if you really need the space savings (like, it's what makes an index fit in RAM) or else we'd see it in use more places. It'd still be kinda neat to actually see those tradeoffs.
- jltsiren 5y agoThere was some interest in the information retrieval research community 10-15 years ago, but I don't think anyone ever found a good application for it. Some limitations of the BWT always got in the way. The BWT sees strings as integer sequences. Either "ABC" and "abc" are two unrelated strings, or you normalize before building the index and lose the ability to distinguish between the two. Search proceeds character-by-character backwards, jumping arbitrarily around the BWT using the same LF-mapping function as when inverting the BWT. You get cache misses for every character. BWT construction is expensive, because you want a single BWT for the entire string collection. There is a ridiculous number of papers on BWT construction, as well as on updating and merging existing BWTs, but the problem has still not been solved adequately. If your data is measured in gigabytes, you can just pay the price and build the index, but a few terabytes seems to be the practical upper limit for the current approaches. You can of course partition the data and build multiple indexes, but then you have to search for each pattern in each index. There is no way to partition the data in a way that different indexes would be responsible for different queries.
- twotwotwo 5y agoAll interesting! Thank you.
- glofish 5y agoAlas the information presented is an over simplification of the process. To actually sequence DNA with this USB thingy you need to prepare a so called sequencing library - and for that you need a fairly well equipped lab - expensive reagents and years of practice and skill ... a mid level biology Ph.D can prepare these ... in addition the flowcell sold by Oxford Nanopore often malfunctions and the whole run is a bust ... (behaves like this since 2014 ... so no, the technology does not seem to improve a whole lot)
- 9tailedkitsune 5y agoYep
- walterbell 5y agoThere are some bio HackerSpace labs with memberships open to the public. London, UK https://biohackspace.org/ https://biohackspace.org/ Brooklyn, NY, https://www.genspace.org/ https://www.genspace.org/ Baltimore, MD, https://bugssonline.org/ https://bugssonline.org/ Australia, https://foundry.bio/ https://foundry.bio/
- unemphysbro 5y agoHappy to see this year. I worked on solid-state nanopore development as a part of my PhD. Now I'm a Data Engineer doing backend work in public sector. :) Here are some press releases related to articles I published during my PhD: https://physics.illinois.edu/news/article/34064 https://physics.illinois.edu/news/article/34064 https://www.sciencedaily.com/releases/2014/10/141014095320.htm https://www.sciencedaily.com/releases/2014/10/141014095320.h...
- throwaway56229 5y ago
- luxpir 5y agoThere is a 3+ year old London-based project, partnered with an established genome sequencing company, doing something highly interesting. They sell swab kits directly, or via NFT purchase, for ~$500 for a 30x near complete sequencing (that's 30 passes for over 99.9% vs 0.2% for 23andme et al). The results are stored in an encrypted AMD SEV-E vault to be accessed by big pharma or individuals, only for specific markers, in exchange for the $GENE token paid directly to the genome owner. Figures touted are $50-80 per request. This token is burned as kits are sold, can be staked, offers rewards like DAO membership, can be gifted to charities researching specific diseases in various populations. It can act as a form of UBI in unbanked populations and puts your DNA back in your control. To me it's the best use of web3 tech I've come across, so disclaimer, I am invested and a DAO member, but it's early in the project still. They are not quite ready for mass marketing. They are moving over to Polygon for very low transaction fees in January, will be launching the first joint NFT/kit sale (the next season might include personal genetically generated art) to fill the vaults with 10k sequenced genomes. They are over half way already through work with charities, but that is the magic number before big pharma can start making queries. Right now though they are quietly building and preparing before marketing plans kick in later in Q1. Take a look at https://genomes.io https://genomes.io where everything is explained in more detail, the team are presented and the tokenomics set out. TL;dr - for $500 right now you can get your entire genome sequenced, stored in a vault to earn you passive income, if you agree to each query. But wait for the NFT vs buying directly, it will have more perks.