5 ms·
I'm the co-founder and CTO at BillionToOne. I'm happy to answer any questions here. I've also posted a slightly more technical explanation of how the test works
by dvdt 6y ago
I'm the co-founder and CTO at BillionToOne. I'm happy to answer any questions here. I've also posted a slightly more technical explanation of how the test works and why it can scale here: https://twitter.com/dtsao/status/1247642005510873088?s=21 https://twitter.com/dtsao/status/1247642005510873088?s=21
Edit: Since our site seems to be overwhelmed at the moment, here's a recap:
We’ve been working hard at BillionToOne on a new COVID-19 test that scales testing to everyone in the US. Our test (1) re-purposes existing infrastructure, (2) eliminates time-consuming RNA extraction, and (3) enables a distributed system for COVID-19 testing.
We need 1 million tests per day to end the stay-at-home orders. Schools are still open in Iceland because they test 15x more than the US does, per capita (https://www.washingtonpost.com/world/2020/04/02/free-coronavirus-test-anyone-this-country-its-possible/ https://www.washingtonpost.com/world/2020/04/02/free-coronav...).
The first thing we figured out is how to run COVID-19 tests on existing automated Sanger sequencers. One sequencer can process up to 3840 samples per day. There are hundreds of sequencers of excess capacity because they were built for the Human Genome Project over 20 years ago.
It would take only 2 sequencers to surpass the current test capacity for all of California. There are far more than 2 sequencers in California (some individual labs have 10 or more).
We tweaked the protocol so COVID-19 could be detected from sequencing data using linear regression. Basically, we add ~100 copies of a known DNA sequence to help us calculate how much virus nucleic acid is in the specimen. It works just as well as gold-standard RT-qPCR.
Lab workflow for COVID-19 testing is traditionally 1. Specimen accessioning, 2. RNA extraction, 3. RT-qPCR 4. Reporting. RNA extraction, in particular, has been a huge bottleneck in terms of reagent shortages and labor-intensiveness.
We showed that we can skip RNA extraction entirely without affecting test sensitivity and limit of detection.
By skipping RNA extraction and using automated Sanger sequencers, we think we can get to an additional 200,000 samples per day test capacity in existing clinical labs.
A distributed system is often the only way to operate at massive scale. A fully distributed system could have different sites and labs responsible for each process and dynamically re-allocate resources based on availability and capacity.
The Broad institute COVID-19 lab has already started doing this. They are asking for specimens to be submitted in a standardized tube format and pre-barcoded. They have essentially distributed the specimen accessioning work.
Because there is a highly developed service industry for Sanger sequencing with <24 hour turnaround, there is an opportunity to further scale up testing by distributing the work to their (currently) idle sequencers.
Distributed testing could scale from 200k to >1 million tests per day, but would require a change in regulations that currently prohibit it.
Thanks to the BillionToOne team for pulling this work together! Next step is to start manufacturing test kits and obtain Emergency Use Authorization from the FDA. We’re eager to work with clinical Lab Directors and contract kit manufacturers.
Edit 2: Link to scientific manuscript: https://www.dropbox.com/s/07esyehsvfpmllc/A%20Highly%20Scalable%20and%20Rapidly%20Deployable%20RNA%20Extraction-Free%20COVID-19%20Assay%20by%20Quantitative%20Sanger%20Sequencing-final.pdf?dl=0 https://www.dropbox.com/s/07esyehsvfpmllc/A%20Highly%20Scala...
- ipsum2 6y agoThanks for the link to the scientific manuscript. Can you speak more about the machine learning aspect? Best of luck on getting FDA approval.
- dvdt 6y ago"When you’re fundraising, it’s AI When you’re hiring, it’s ML When you’re implementing, it’s linear regression" The core of our machine learning is Ax=b :grins: More seriously, the main reason why traditional sanger sequencing can't be used for COVID-19 testing is because it would be unclear whether a lack of signal is truly due to lack of virus, or if it is just because the assay failed (happens all the time!) What we've done is introduce a reference sequencing signal that is biochemically very similar to viral RNA, but produces a distinct vector of electrical signals that is different from the signals emitted by viral RNA. Since we know what both the reference and viral signals look like, we can perform linear regression analysis to fit the linear combination of viral and reference signals that best match our data.
- dang 6y agoFor HN it's better to drop the fundraising language and use the implementing language, so I've s/machine learning/linear regression/'d your text above.