13 ms·
We're wasting money by only supporting gzip for raw DNA files
- zellyn 4y agoAre these files already just diffs from baseline/standard full DNA scans, or are they including all the data? Seems like most scans will only differ a little…
- biotinker 4y agoThese are FASTQ files[0], which are the actual raw bases read by the sequencer, and a quality score for each base. This is upstream of comparison to a known genome any other processing. You're correct that there's a lot of redundancy in these files, not least because each area of the genome is generally covered several times over. Removing this redundancy (or using it for error correction) is a big part of the downstream processing of this file. [0]: https://en.wikipedia.org/wiki/FASTQ_format https://en.wikipedia.org/wiki/FASTQ_format
- cjbgkagh 4y agoThese snippets can be used to reconstructing the whole genome sequence but it's done statistically so they're not 100% confident. Some sections have more coverage than others. You want to keep the raw reads around just in case. Longer higher quality reads have been developed and for them less coverage will be needed and this will produce smaller files with higher degrees of confidence.
- biotinker 4y agoDepends on the use case. If one is trying to broadly sequence a genome or exome, all of that is correct. For something like a liquid biopsy test that's looking for circulating tumor DNA, a given mutation may only exist in 0.5% of the reads that cover a particular locus. Then you have to deduplicate to tell which are PCR duplicates of the same original strand, and which have different progenitor fragments. Low-coverage long reads are bad for this.
- cjbgkagh 4y agoI had not thought of that use case but that does make sense.
- otherme123 4y agoThe whole point behind Next Generation Sequencing is to get a lot of redundant and cheap raw data, but with some errors. This way you sequence the same base of the genome say 10 times or more. If 9 of them are the same as the reference, and 1 is another base (and this base is low quality), the Variant Caller can be confident that was a sequencing error, and the position is the same base as the reference. But if you have 5 reference and 5 of other base, it's probably heterozigous for that position. If you have more of other base, is homozigous for that other base.
- deleted 4y ago[deleted]
- dekhn 4y agoThey are FASTQ files, so they are the reads that came off the sequencer- plus metadata and error scores- represented as ASCII A, T, G, C. This isn't a diff from a reference genome, it's a ton of reads (files often 30+GB). It's really only useful as an archive format, and even then, it's a pretty bad one for many reasons. An interesting thing about the file format is that it's three lines in a row each with different coding properties, so if you break the data into three streams and encode each one on its own, the overall compression is somewhat better, because the dictionary and lookback don't have to accomodate the metadata, DNA, and error lines together, giving better dictionary and lookback hit rates. These formats are terrible for downstream analysis: they are typically not indexed or sorted in a useful way, so it's common to convert them to other formats (BAM, CRAM), which are optimized for compressing DNA data (and likely exceed zstd in some circumstances).
- otherme123 4y ago> These formats are terrible for downstream analysis: they are typically not indexed or sorted in a useful way, so it's common to convert them to other formats (BAM, CRAM), which are optimized for compressing DNA data (and likely exceed zstd in some circumstances). FASTQ are raw data reads. BAM and CRAM are reads mapped against a genome.
- biotinker 4y agoBAM/CRAM reads are often mapped against a genome, but they needn't be. It's not uncommon in industry to convert FASTQ to unmapped BAM in order to only have a single file format, improve compression, etc.
- zellyn 4y agoThanks all for the interesting and useful replies!
- xbar 4y agozstd -10 is so very fast and so much better than gz that I am surprised every time I find someone still using gz for large files. When I can get away with -16 for large file long term storage, I use it.
- warmwaffles 4y agoBecause `tar -czvf` is muscle memory at this point for many.
- croemer 4y agoYep, I've been advocating zstd since I started working in the field. It compresses and decompresses so much faster than xz and is much much more compact than gzip. All public SARS-CoV-2 consensus sequences are ~300GB uncompressed. If you compress using gzip you end up with ~30GB. With xz/zstd you can get it down to ~2GB. However xz takes ~40min to uncompress, whereas zstd can do it in ~8min.
- mort96 4y ago[flagged]
- Underphil 4y agoWeird take seeing as it's licensed under GPL.
- mort96 4y ago[flagged]
- chungy 4y agoYou're free to march on with your weird bias. The rest of us can benefit.
- mort96 4y agoCan't "the rest of us" make other good compression formats?
- mikepurvis 4y agoBecause large companies are willing to pay people market salaries to work on this stuff, and then give it away? Isn't that what we want, a model where free software development is properly compensated?
- wpietri 4y agoGlad to see zstd getting some love here. At a previous job I needed to capture and store large amounts of JSON about social media. I did an extensive compression bake-off taking into account both compression and storage costs, and for us the winner was zstd with the compression dial turned up a fair bit. (Sorry I don't have numbers here, but they're all in the hands of the previous employer. But I think we settled on -13 as cost-optimal if we were storing data on AWS for a year.)
- Retr0id 4y agoI'm unfamiliar with the bioinformatics scene, but I can imagine there's a lot of "legacy" hardware and software lying around that can't easily be updated to support superior formats. Perhaps an interim solution would be something like "zstd+metadata", where the metadata is sufficient to transparently and efficiently reconstruct a gzip-compressed file on-demand. (similarly to how JPEG-XL allows oldschool JPEGs to be recompressed without losing any data) This could have a bit of compute overhead, but since gzip decompression is so single-threaded I think you could do the conversion in parallel without actually slowing down the hot path. So, the performance would be approximately the same as using gzip (assuming decompression is compute-bound, not io-bound), but with all the storage benefits of using zstd.
- croemer 4y agoYeah that's exactly the issue. A lot of tools are written once and rarely updated, they're considered "done". And they only support gz out of the box because they were written 10 years ago.
- jltsiren 4y agoI think the tools people actually use are rarely that old. Almost all new tools support gzip by default, because almost all public data is gzip-compressed. Developers often don't bother adding zstd support, because there is little zstd-compressed data around. Because the labs developing new tools are rarely the ones that have to store petabytes of data, they don't benefit from zstd directly. And once the active development is over, most tools receive only minimal support, because nobody wants to pay for maintaining the long tail of average tools. Also, zstd is the latest reason why gzip is still widely used. Every time someone releases a new general-purpose data compressor that becomes popular, gzip gains a few more years. Because there have been too many gzip replacements over the years, none of them has managed to become popular enough to actually replace gzip before the attention moves to the next state-of-the art compressor. If people agreed this time that zstd is what everyone will be using for the next 20 years, bioinformatics would probably move to it in a few years.
- asdff 4y ago
- wheresmycraisin 4y agoA much bigger problem than the compression format is the fact that biologists over-sequence. You don't need 500x coverage, trust me, please learn to design your experiments.
- chrisamiller 4y agoThat's a pretty expansive statement. Are you doing error-corrected sequencing to look for rare mosaic variants? Do you care about transcripts expressed at very low levels? Are you trying to tease apart subclonal variation in cancer? All of these could be good reasons to sequence deeply. If anything, I see the opposite problem much more frequently: People who cheap out on the sequencing depth, neutering their statistical power!
- biotinker 4y ago> You don't need 500x coverage, trust me, please learn to design your experiments Maybe you don't need 500x coverage, and if you don't, then I'm glad you're not sequencing to that depth. But please don't tell other people they're doing something wrong when you don't understand their use cases. Making the blanket statement that 500x coverage is generally not needed for any experiment is patently false. I've built clinical liquid biopsy tests where it was a QC failure if particular sites of interest were under 5000x coverage. Sequencing such sites to only 500x coverage would be a waste of everybody's time, because it would be impossible to gain meaningful results from so little data.
- pneumic 4y agoSome people do need it for R&D.
- firecraker 4y agoTo those who don't know.. file formats in genetics is already a big mess.
- otherme123 4y agoTrue. And a huge share of formats are just re-inventing the wheel of relational databases but in plain text.
- asdff 4y agoI would assume most datasets people are using aren't large enough to justify using a database versus parsing a flat file. You also have to realize this is a field of scientific code, where it matters more that you spend less time coding and more time interpreting results, versus spending time optimizing the pipeline to minimize compute time, and you might be working on a university cluster where your compute is powerful and quite cheap. For people who might work on clinical pipelines that will continue to be reran time again over a vast growing amount of patient data, they probably already put their data into databases. For your academic post doc working on 2000 samples from an experiment for one paper before they find another job doing something else entirely in two years, a flat file is fine.
- otherme123 4y agoI was thinking for example in VCF files. A metadata header, a main table with eight clear columns and a ninth column that works as a "put here whatever you need", and then the related data for each sample in extra columns. Next thing you have is a set of tools to recreate a small subset of SQL, to index the file, to add in bulk, to edit the metadata... The typical VCF has data enough to be a SQLite, and nobody parses the VCF directly but with tools. This ends in a sad number of bio-scientists that cannot do the simplest SQL query, but know perfectly vcftools, samtools, bedtools and others (or have them hardcoded in shell scripts). Those formats start so simple you can "parse" them with grep, cut, wc and paste, but soon they need special tooling and get feature creep.
- transcriptase 4y ago
- arminiusreturns 4y agoAs someone who helped build a fastq data processing flow for a genetics company during the sequencing price drop era, iirc we had issues with data corruption with some of the other formats in tests. Funny to think, hey, I got my first taste of very large data management because of this!
- Scaevolus 4y agoThere are compressors specialized for FASTQ that are faster and denser than zstd. FASTQ is the the most common format for storing DNA sequencing data-- a text file including metadata, sequence, confidence scores, etc. http://kirr.dyndns.org/sequence-compression-benchmark/ http://kirr.dyndns.org/sequence-compression-benchmark/ fastqz (not the best name!) supports reference compression too, for smaller files when they're reads compared to a reference genome. Here's an 2013 paper considering the problem: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0059190 https://journals.plos.org/plosone/article?id=10.1371/journal... Still, zstd is widely available and a simple drop-in replacement for gzip.
- _yb2s 4y agoReference compression is interesting... technically, this type of compression should be possible without a reference. E.g. in a worst case scenario, you do a genome assembly from the reads, and then compress against that reference. Fascinating to think that in principle, genome assembly and fastq file compression have many parallels, and are essentially the same problem. For example, once you assemble the genome, you can just track the coordinates and 'diff' for each read and the original FASTQ file can be regenerated exactly.
- bnprks 4y agoOne recent method that does this is SPRING: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6662292/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6662292/ Although I'm not familiar with all the details of the method, my understanding is that the reference used by compression can be extremely low quality for biological purposes while still providing good compression -- making it possible to make a quick-and-dirty reference with much less computational time.
- v8xi 4y agoThat would be true for FASTA compression but FASTQ files contain other data. With a MiSeq, for example, 1/3 of the data is the DNA sequences, 1/3 is a quality score (0-37) and 1/3 is a header which includes things like flow cell ID but also a lot of info on the physical registration address of the cluster on the flow cell. In some applications this info is not required, but is from a scientific/data integrity perspective
- evrydayhustling 4y agoWho would implement this suggestion? Is it an appeal to folks at large writing tools that interact with genomics data? Due to higher decompression cost, the opportunity here seems localized to long-term storage. It feels like it would make more sense as a project (or product!) that implemented an efficient long-term archive (perhaps with a less compressed LRU cache in front).
- michaelbarton 4y agoYes exactly. If the most common tools supported zstd then it would make it easier to store zstd FASTQ. At the moment I believe almost none of them do.
- asdff 4y agoFor those toolmakers they would probably say something like "our tool supports uncompressed data specifically so you can add whatever compression you want in your pipeline for your use cases."
- chrchang523 4y agoIt addresses different bioinformatic use cases, but plink2 is an open-source tool that is widely used and makes extensive use of zstd. In particular, it accepts both gzip- and zstd-compressed text files as input, but output commands that can generate large text files often only have a zstd compression option.
- qclibre22 4y agogzip and xz are usually installed these days, zstd isn't.
- retrac 4y agoUbuntu, Fedora, and Arch use zstd for their packages, all switching in the last few years. FreeBSD seems to have it in base. OpenBSD doesn't have a man page for it. I think Mac OS doesn't include it.
- asdff 4y agoClusters might be running distros like centos though.
- croemer 4y agozstd will be installed soon, it's easy to install anyways
- charonn0 4y agoI don't know anything about DNA files, but I wonder how gzip (deflate) would fare if one of the non-default compression strategies were used. e.g. the RLE strategy is good for data with many compressible patterns, while the filtered or Huffman-only strategies are good for data with few compressible patterns. The default strategy is a balance between the two extremes, which isn't tuned for a specific kind of input.
- kxc42 4y agoI'm not sure why gzip still pops up for FASTQ data, as it is quite easy to bin the quality scores, align it against a reference genome and compress it as e.g. CRAM [1,2]. With 8 bins, the variant calling accuraccy seems to be preserved, while drastically reducing the file size. [1]: https://en.wikipedia.org/wiki/CRAM_%28file_format%29 https://en.wikipedia.org/wiki/CRAM_%28file_format%29 [2]: https://lh3.github.io/2020/05/25/format-quality-binning-and-file-sizes https://lh3.github.io/2020/05/25/format-quality-binning-and-...
- jefftk 4y agoYou don't necessarily have a reference genome to align to. For example, I've recently been working with wastewater metagenomics where (a) the sample consists of a very large number of organisms and (b) we don't have reference genomes for most of these organisms anyway.
- kxc42 4y agoThat can be a challenge, but you can also build an "artificial" reference genome. You just use it for compression, not for any real analyses. This would allow you to still use alignment-based compression. But I agree with you: it really depends on the type of the data.
- dekhn 4y agoIt would be nice also that the artificial reference represented global population structure- for example, the larger the genetic distance between an individual who is sequenced, and the identity of the person who makes up the reference (an amalgam of several individuals from a common US population), the less compression you get. Instead, it seems like you could create the "genome that is the shortest distance to all other genomes" (a centroid of cluster centroids) and then the standard deviation of your compressed sizes should be much smaller.
- v8xi 4y agoWell I think the issue with wastewater and other screening tech is that there is no global average reference genome. In that case they're sequencing everything from phages, viruses (human and plant), bacteria, fungi, plants/animals and human...its an everything soup.
- dekhn 4y agoI have some questions for folks who are working in this field. In particular, are you holding onto your FASTQs for a long time (> 1 year), and if so, why? Is max compression, or ease of analysis more important (IE, do you access the data during the retention period, do you ETL it to another format, etc). How much of your budget represents storage costs which could be affected by compression? I'm curious because in the past I've seen people are very cost sensitive but also want to keep lots of data that seems to go mostly unused.
- transcriptase 4y agoUsually fastQ filed that aren’t publicly available are retained pretty much indefinitely. The reason is, in my experience, the ability to align the sequence to new and improved genome assemblies or using/benchmarking new tools altogether. I’ve had various reasons to go back to raw sequences that had been stored for 8+ years to reanalyze them alongside newly generated data.
- dekhn 4y agoIs the cost worth it? Many groups are dealing with literally petabytes of fastq.gz and they usually sit around un-re-analyzed (I keep my own personal genome BAMs around and occasionally un-map them and re-map them to new references).
- transcriptase 4y agoIt really depends on the population or potential future use-case. I’ve known many labs where cost wasn’t a concern because their institution was associated with (or had their own) cluster that was under-utilized. I think there’s also something to the fact that the cost of storing sequence pales in comparison to the sample collection, DNA extraction, library prep, and sequencing itself that if you can afford to generate that much sequence then storing it isn’t much of an issue. And if you really wanted to cut costs it’s easy to upload to NCBI rather than deleting it. I also think there’s a certain degree of fear in not being able to generate the data again if you needed to, due to lack of funding or the organism itself not being available.
- retrac 4y agoGenome storage and compression is interesting. Most long sequences have a lot of internal redundancy/repetition that can be effectively compressed by standard algorithms, but something tuned specifically to the task, can do better. Then there is the question of storing many sequences. Whole genome sequences of thousands of bacteria from an experiment, for example. Each is almost identical to the other. This benefits from an algorithm that can compress that efficiently in terms of space, but also allows any part to be recalled random-access reasonably efficiently in terms of time. It'd also be nice to be able to add more genes to the database without having to re-compress the entire thing. Wikipedia has a page on the topic: https://en.wikipedia.org/wiki/Compression_of_genomic_sequencing_data https://en.wikipedia.org/wiki/Compression_of_genomic_sequenc...
- sl-dolt 4y agoYou can stream gzipped files to work with them out of memory. Can you stream this? Is that even a use-case to consider?
- rabernat 4y agoThe Zarr format is used in some genomics workflows (see https://github.com/zarr-developers/community/issues/19 https://github.com/zarr-developers/community/issues/19) and supports a wide range of modern compressors (e.g. Zstd, Zlib, BZ2, LZMA, ZFPY, Blosc, as well as many filters.)
- dekhn 4y agoI would zarr for dense matrices, mostly (I use them with microscope images). I see it also used for frequency/spatial observations in genomic imaging. But I prefer parquet for most direct analysis of sequence, since it's the format best integrated with big data analytics. I care much less about total compression size than I do the ability to decompress the data I need quickly (say, to ETL it to a featurization pipeline).
- eternalban 4y agoI grep'd for parquet and yours is the only comment that mentions it. parquet -> arrow -> AI
- dekhn 4y agoSee https://techcommunity.microsoft.com/t5/healthcare-and-life-sciences/genomic-data-in-parquet-format-on-azure/ba-p/3150554 https://techcommunity.microsoft.com/t5/healthcare-and-life-s... and https://adam.readthedocs.io/en/latest/api/genomicDataset/ https://adam.readthedocs.io/en/latest/api/genomicDataset/ as well as https://hail.is/ https://hail.is/
- eternalban 4y agoMany thanks. Q: what's the story around versioning, provenance, reproducibility, (etc.ops) in your domain? I've seen various bolt x on/along git variants. Wondering if its worth the effort to make something to address that.
- riverdweller 4y agoA few years ago I built some tools https://github.com/tf318/tamtools https://github.com/tf318/tamtools to store alignments against two different reference assemblies in an efficient way (taking advantage of the fact that the majority of each alignment to different assemblies would in fact be the same, just shifted in position). The intent was to enhance this to store alignments against multiple references as new references are published, and probably to rewrite in Rust or C rather than the initial Python version. In retrospect I would be interested to know whether this domain-specific compression effort, with zstd to the resulting "hybrid" alignment, would be more efficient than just letting zstd do its own thing with a full set of individual alignments against the different references.
- n4jm4 4y agoRewrite our DNA from quadature to base64. Much more efficient.
- rwmj 4y agoI'm slightly surprised that speed and file size would be the only considerations. For an archival format wouldn't you at least want a format with very robust error detection, even error correction? None of the proposed formats have this, basically just having a simple 4 byte CRC or XXHASH. Some kind of random access to the compressed file might be useful too (which xz has).
- phantop 4y agoI wonder if making using of zstd's long mode (`--long`) and/or multithreading support (`-T0` to use all threads) would close the displayed performance gap with `pigz`. Doesn't seem like either were used, which is odd considering the comparison made and the file used.
- michaelbarton 4y agoI did play with long mode and it indeed works much better. I omitted it for brevity and to drive faster the main point about gzip. Pretrained dictionaries could also be useful too.
- polyrand 4y agoOne thing that's missing from the article when comparing to `pigz` is that you can use the `-T0` flag in `zstd` and it will parallelize according to the number of CPUs. On some limited benchmarks, I found it to be much faster and worth using. Some `zstd` installations come with `zstdmt`[0] as an alias to `zstd -T0`. [0]: https://manpages.debian.org/bullseye/zstd/zstdmt.1.en.html https://manpages.debian.org/bullseye/zstd/zstdmt.1.en.html
- GistNoesis 4y agoI'm not in the bioinformatic domain but maybe there are some legacy tools that depend on some specifics of gzip. I was thinking something like the possibility to seek through a compressed file (after having built an index). Your DNA file is probably stored on a shared folder somewhere (if it's big you probably don't want to copy it to every workstation) and you point your software to the gz file directly, it create a seek-index on the first use, and then when you need view only a small section of the file you don't have to extract the whole file. Some more advanced compression like zstd might use some form of dictionary to do the compression/decompression, which may be bigger (up to 32MB? max dictionary size for zstd whereas lz77 used by gzip use a dictionary based on a sliding window of the last 32K token) which you may have to transfer.
- v8xi 4y agoCan confirm that the data is often, though not always, stored in a shared directory and that yes, many tools can/do read gzip files directly which probably hinders the adoption of better compression algorithms.
- asdff 4y agoIf they do then they are just probably running zcat for you for convenience sake. You can always incorporate better compression by adding it to your pipeline, and piping in uncompressed data to the legacy tool then compressing the resulting output with whatever tool you want.
- daemonk 4y agoThe best storage format for genomics data is DNA. At some point in the future, it might actually be cheaper to just re-sequence than to store the fastq. Instead of optimizing the digital infrastructure, it might be better to just optimize the lab infrastructure for storing the physical DNA.
- postalrat 4y agoIs it the best? It's not compressed.
- elmolino89 4y agoCompression of FASTQ files can be greatly improved by sorting/clustering the reads. I use clumpify from BBMap for that. The bad: clumpify does not support zstd at this point.
- borland 4y agoThis seems silly. Any company that is in the business of storing tons and tons of DNA data will probably be recompressing it for archive purposes using better algorithms already. If they're not, their loss. For smaller players where maximum tool interoperability is important, gzip seems good enough.
- wolf550e 4y agozstd has a long range mode, which lets it find redundancies a gigabyte away. Try --long and --long=31 for very long range mode. zstd has delta / patch mode, which creates a file that stores the "patch" to create a new file from an old (reference) file. See https://github.com/facebook/zstd/wiki/Zstandard-as-a-patching-engine https://github.com/facebook/zstd/wiki/Zstandard-as-a-patchin... See the man page: https://github.com/facebook/zstd/blob/dev/programs/zstd.1.md https://github.com/facebook/zstd/blob/dev/programs/zstd.1.md
- pinetroey 4y agoAre there any machines that can home sequence(machine @ home) for a resonable cost?
- mfld 4y agoThe best fit is probaly the Oxford Nanopore MinION. However, it would still be good to have a range of lab equipment available.
- cornutopia2 4y agoMany DNA specialized formats are delta from a "baseline" model to the actually described DNA strand. `zstd` has a `--patch-from` mode, which does essentially the same thing, computing a new file using another file (typically an older version) as a reference. When similarities are larges, it leads to a huge reduction in compressed content. I wonder what would be the performance of `zstd --patch-from=` when employing the same reference as these specialized DNA compressors.
- mfld 4y agoIf you have re-sequencing data of model species (which applies to >80% of generated sequencing data), the storage issue is often solved using CRAM/BAM formats. The FASTQ can be reconstructed if unmapped reads are stored in the file. More general (pre-alignment) sequence compression methods never really took of (e.g. https://github.com/BEETL/BEETL https://github.com/BEETL/BEETL). Probably because it helps so much to have common format that most workflows can start with. Here, the replacement of gzip with zstd would be a lower hanging fruit.
- jbverschoor 4y ago> The ztsd -15 command takes ~70s which is 50% longer than pigz -9 at ~35s. No, it’s 100% longer, twice as long, and the inverse is 50%
- michaelbarton 4y agoGood catch
- AbrahamParangi 4y agoIs there any good work on NN based compression for DNA? For those not aware one can, in general, use a predictive model to compress data using arithmetic coding.