4 ms·
Usually fastQ filed that aren’t publicly available are retained pretty much indefinitely. The reason is, in my experience, the ability to align the sequence to
by transcriptase 4y ago
Usually fastQ filed that aren’t publicly available are retained pretty much indefinitely. The reason is, in my experience, the ability to align the sequence to new and improved genome assemblies or using/benchmarking new tools altogether. I’ve had various reasons to go back to raw sequences that had been stored for 8+ years to reanalyze them alongside newly generated data.
- dekhn 4y agoIs the cost worth it? Many groups are dealing with literally petabytes of fastq.gz and they usually sit around un-re-analyzed (I keep my own personal genome BAMs around and occasionally un-map them and re-map them to new references).
- transcriptase 4y agoIt really depends on the population or potential future use-case. I’ve known many labs where cost wasn’t a concern because their institution was associated with (or had their own) cluster that was under-utilized. I think there’s also something to the fact that the cost of storing sequence pales in comparison to the sample collection, DNA extraction, library prep, and sequencing itself that if you can afford to generate that much sequence then storing it isn’t much of an issue. And if you really wanted to cut costs it’s easy to upload to NCBI rather than deleting it. I also think there’s a certain degree of fear in not being able to generate the data again if you needed to, due to lack of funding or the organism itself not being available.
- bnprks 4y agoOne useful perspective to consider is the costs of computation compared to the costs of the experiment. Considering sequencing costs alone (ignoring costs of obtaining & processing a biological sample), Illumina's current cheapest sequencing kits - Novaseq S4 300 cycles - cost ~$15k for 3000 Gbases of data (whole-genome sequencing for about 24 people). As gziped fastqs, that data will occupy about 1.3TB of space. Extremely high-volume purchasers may be able knock that sequencing price down some, but the storage challenge starts to seem less intimidating when you realize spending 5% of your sequencing budget on storage would give you a budget of $750/TB of data.
- asdff 4y agoFastq file will be a little smaller than a bam file
- chrisamiller 4y agoIn most cases* the raw data can be recreated from the aligned BAM/CRAM file, so it's often sensible to toss the raw FASTQs after alignment. But yeah, I go back to 10+ year old genome sequences more than some might expect. *In the absence of read trimming or discarding unmapped reads