3 ms·
plus a special sauce for counting the number of specific bp repeats, due to in-del events, this is not something I am not too familiar, but presumably the numbe
by jrm5100 10y ago
plus a special sauce for counting the number of specific bp repeats, due to in-del events, this is not something I am not too familiar, but presumably the number of a specific k-mer repeats you have in these genes of interest might correlate to a specific type of cancer? (would love to hear someone who is an expert in this field their opinion).
"Copy number variant" refers to larger deletions and duplications that can occur in the genome. There isn't some specific cutoff for size, but some examples in these kinds of genes would be an entire exon or gene. There are countless studies that find correlations between specific variants or CNVs and risk of cancers.
Standard variant detection is pretty straightforward. CNVs are harder because they are longer (several hundred to several thousand base pairs) than the raw data (150 to 250 bp for Illumina)- you don't get single reads that span the entire variant. You have to normalize then look for differences in coverage, or look for split reads (where the read is aligned on the border of one of these CNVs).
This kind of funding baffles me because they don't seem to be proposing anything new at all (maybe slightly better CNV detection?) and there are already lots of labs/companies doing this kind of testing. Maybe they are working on being very efficient to offer a better price.
- noname123 10y agoThanks for your detailed explanation. Just out of curiosity and to follow-up, presumably this is a example of the list of detected CNVs in a TCGA Breast Cancer data-set you're referring to: http://cancer.sanger.ac.uk/cosmic/gene/analysis?ln=BRCA1#cnv_t http://cancer.sanger.ac.uk/cosmic/gene/analysis?ln=BRCA1#cnv... According to Sanger (or maybe TCGA?), a gain is when a genomic region (for a diploid) has more than five absolute copies of this region and a loss is when the genomic region has no reads ((http://cancer.sanger.ac.uk/cosmic/help/cnv/overview http://cancer.sanger.ac.uk/cosmic/help/cnv/overview), where the copy number is perhaps determined by that normalized distribution of read coverage across the reference genome? (http://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-14-S11-S1#Fig1 http://bmcbioinformatics.biomedcentral.com/articles/10.1186/...). This is for CNVs that are longer than the 150-200bp Illumina fragments (Fig1c. Read Depth method, e.g., exome#3 looks like it has two absolute copies vs exome #1 and #2) Then for small CNVs that perhaps span that 150-200bp fragment, we use the split read method to filter for incompletely mapped reads that are only aligned on the edges to the reference. This implies that there was a duplication event that expanded that sequence? (Fig 1b. Split Read method). Presumably, the pipeline would determine the CNV sites in a specific patient sample, then cross-reference with the TCGA CNV data-set and come up with correlation score of how much those CNVs sites match with consensus CNVs in the cancer data-set? Thanks again for your detailed breakdown.
- jrm5100 10y agoThe Sanger/TCGA (The Cancer Genome Atlas) stuff seems to be specific to microarray data which is different (older, more expensive) than the newer high-throughput data. The figure you linked is a good explanation. The split read method is helpful for finding the edges of the CNV, while the number of reads (relative to other regions that were tested) can give an idea of the number of copies. The problem is that these methods all have their own unique biases/noise that makes it non-trivial to figure out the absolute copy number change. Ideally they would find a similar CNV that has some clinical association. The DGV has a lot of reference CNVs. Here are some in BRCA1: http://dgv.tcag.ca/gb2/gbrowse/dgv2_hg19/?name=id:3087443;dbid=gene:database http://dgv.tcag.ca/gb2/gbrowse/dgv2_hg19/?name=id:3087443;db...
- noname123 10y agoThanks jrm5100 for the link. I see the variants under the "DGV Structural Variants" track. Really appreciate your explaining what CNVs are and also following up on my questions/confusions!