16 ms·
Xz format inadequate for long-term archiving (2017)
- nurettin 8y agoToo bad for arch https://www.archlinux.org/news/switching-to-xz-compression-for-new-packages/ https://www.archlinux.org/news/switching-to-xz-compression-f...
- saghm 8y agoIs this really an issue for this use case? My naive take is that since Arch updates packages so often, "long-term storage" doesn't come up that much in practice.
- mikepurvis 8y agoDefault compression for debian packages is xz as well: http://manpages.ubuntu.com/manpages/xenial/en/man1/dpkg-deb.1.html http://manpages.ubuntu.com/manpages/xenial/en/man1/dpkg-deb....
- deleted 8y ago[deleted]
- agumonkey 8y agothis is from 2010, I guess if xz was bad for this use case they'd know by now
- cpburns2009 8y agoIt may not be a good choice for long-term data storage, but I disagree that it should not be used for data sharing or software distribution. Different use cases have different needs. If you need long-term storage, it's better to avoid lossless compression that can break after minor corruption. You should also be storing parity/ECC data (I don't recall the subtle difference). If you only need short to moderate term storage, the best compression ratio is likely optimal. Keep a spare backup just in case.
- planteen 8y agoI've used XZ to compress tarballs of backup. XZ was useful so I could store more backups on an external hard drive. I have seen bit rot on some of these files (stored on a magnetic HDD), in the sense that the md5sum of the .tar.xz archive no longer matches when it was created. What do you suggest for creating parity/ECC in this case? I'm aware of parchive, but is that the right choice and in what configuration?
- cpburns2009 8y agoKeep in mind I'm not an archival expert so you should do your own research. That being said, currently I'm using pyFileFixity [1] to generate the hashes and ECC data for my personal backups. I write them to M-Disc Blu-rays using Dvdisaster [2] which can also write additional ECC data. After a lot of googling and reading this useful Super User question [3], and this extensive answer [4] I settled on this setup. I must admit that I am guilty of storing images as JPGs and compressing most most of my files in ZIPs for convenience. [1]: https://github.com/lrq3000/pyFileFixity https://github.com/lrq3000/pyFileFixity [2]: http://dvdisaster.net/en/index.html http://dvdisaster.net/en/index.html [3]: https://superuser.com/q/374609/52739 https://superuser.com/q/374609/52739 [4]: https://superuser.com/a/873260/52739 https://superuser.com/a/873260/52739
- zokier 8y agoThe whole structural adaptive encoding seems like massive overcomplication. I feel like clever tricks such as that serve only to bite in the ass when you need it the most. Same goes for the bit jpeg. Sure, it might not be ideal technically, but recommending JPEG2000 (presumably as there is no JPEG2) with its ridiculously poor software support seems weak too. What use is robust file that you can't open?
- Lionsion 8y agoWhat are better file formats for long term archiving? Were any of them designed specifically with that use case in mind?
- cpburns2009 8y agoThere's a post on Super User that contains useful information: "What medium should be used for long term, high volume, data storage (archival)?" https://superuser.com/q/374609/52739 https://superuser.com/q/374609/52739 It mostly focuses on the media instead of formats though.
- Lionsion 8y agoThanks, I'll take a look. Though I think I have the media question answered, and I settled on M-DISC for personal stuff (https://en.wikipedia.org/wiki/M-DISC https://en.wikipedia.org/wiki/M-DISC). It only has special requirements for writing, reading can be done on standard drives.
- cpburns2009 8y agoI went with M-Disc too and an LG Blu-ray burner. I think you only need a special burner if you're using the DVDs. I want to say most Blu-ray burners work.
- paulmd 8y agoIt all depends on what your definition of "high-volume" is, and just how "archival" your access patterns really are. Amazon Glacier runs on BDXL disc libraries (like a tape library). There's nothing truly expensive about producing BDXL media, there just isn't enough volume in the consumer market to make it worthwhile. If you contract directly with suppliers for a few million discs at a time, that's not an issue (you did say high-volume, right?). https://storagemojo.com/2014/04/25/amazons-glacier-secret-bdxl/ https://storagemojo.com/2014/04/25/amazons-glacier-secret-bd... For medium-scale users, tape libraries are still the way to go. You can have petabytes of near-line storage in a rack. Storage conditions are not really a concern in a datacenter, which is where they should live. (CERN has about 200 petabytes of tapes for their long-term storage.) https://home.cern/about/updates/2017/07/cern-data-centre-passes-200-petabyte-milestone https://home.cern/about/updates/2017/07/cern-data-centre-pas... If you mean "high-volume for a small business", probably also tapes, or BD discs with 20% parity encoding to guard against bitrot. Small users should also consider dumping it in Glacier as a fallback - make it Amazon's problem. If you have a significant stream of data it'll get expensive over time, but if it's business-critical data then you don't really have a choice, do you?
- doubledad222 8y agoThank you for sharing this. I am in charge of archiving the family files - pictures, video, art projects, email. I want it available through the aging of standards and protected against the bitrot of aging hard drives. I'll be converting any xz archives I get into a better format.
- Skunkleton 8y agoYou should also write out ECC information.
- moviuro 8y agoMix and match, according to criticity and max affordable data loss: multiple locations, multiple solutions, multiple local copies (e.g. one cloud solution + DVD + NAS). See: https://www.backblaze.com/blog/the-3-2-1-backup-strategy/ https://www.backblaze.com/blog/the-3-2-1-backup-strategy/
- arundelo 8y agoI upvoted this because it seems to make some good points and I think the topic is interesting and important, but I can't understand why the "Then, why some free software projects use xz?" section does not mention xz's main selling point of being better than other commonly used alternatives at compressing things to smaller sizes. https://www.rootusers.com/gzip-vs-bzip2-vs-xz-performance-comparison/ https://www.rootusers.com/gzip-vs-bzip2-vs-xz-performance-co...
- wyldfire 8y ago> compressing things to smaller sizes. ...relative to ... ? Is it better than lzip? lzip sounds like it would also use LZMA-based compression, right? This [1] sounds like an interesting and more detailed/up-to-date comparison. Also by the same author BTW. [1] https://www.nongnu.org/lzip/lzip_benchmark.html#xz https://www.nongnu.org/lzip/lzip_benchmark.html#xz
- derefr 8y agoRelative to the compression formats people were aware of at the time (which didn't include lzip.) People began using xz because mostly because they (e.g. distro maintainers like Debian) had started seeing 7z files floating around, thought they were cool, and so wanted a format that did what 7z did but was an open standard rather than being dictated by some company. xz was that format, so they leapt on it. As it turns out, lzip had already been around for a year (though I'm not sure in what state of usability) before the xz project was started, but the people who created xz weren't looking for something that compressed better, they were looking for something that compressed better like 7z, and xz is that. (Meanwhile, what 7z/xz is actually better at, AFAIK, is long-range identical-run deduplication; this is what makes it the tool of choice in the video-game archival community for making archives of every variation of a ROM file. Stick 100 slight variations of a 5MB file together into one .7z (or .tar.xz) file, and they'll compress down to roughly 1.2x the size of a single variant of the file.)
- LeoPanthera 8y agoCan you provide an example of such a .xz file?
- kazinator 8y ago> The xz format lacks a version number field. The only reliable way of knowing if a given version of a xz decompressor can decompress a given file is by trial and error. Wow ... that is inexcusably idiotic. Whoever designed that shouldn't be programming. Out of professional disdain, I pledge never to use this garbage.
- zzzcpan 8y agoWelcome to the world of software I guess. Non idiotic things are rare here.
- kazinator 8y agoNot w.r.t. that level of idiotic; and in FOSS, at least, we should be able to eject the idiotic. Thanks in part to articles like this, we can.
- menacingly 8y agoHistrionic reactions don't improve the overall quality of software. We certainly should have environments where we can tell someone code is shit, it's just silly and counterproductive to then leap to attacks on the abilities on the person behind it.
- kazinator 8y agoImproving badly designed software that is unnecessary in the first place is foolish; just "rm -rf" and never give it another thought.
- davidw 8y agoFWIW, xz is also a memory hog with the default settings. I inherited an embedded system that attempts to compress and send some logs, using xz, and if they're big enough, it blows up because of memory exhaustion.
- pmoriarty 8y ago"xz is also a memory hog with the default settings" Then why use the default settings? I tend to use the maximum settings, which are much more of a memory hog, but I have enough memory where that's not an issue. Just use the settings that are right for you.
- davidw 8y agoYou'd have to ask the guy who wrote the code in the first place. I think he saw "'best' compression" and stopped looking there.
- pmoriarty 8y agoI didn't mean to ask why the defaults are defaults, but rather why anyone would use the defaults rather than settings more appropritate to their use case? It's not like xz is unable to be lighter on memory, if that's what you want. It's an option setting away.
- davidw 8y agoTo clarify: you'd have to ask the guy who wrote our code.
- tedunangst 8y agoAre these concerns, about error recovery, outdated? If I want to recover a corrupted file, I find another copy. I don't fiddle with the internal length field to fix framing issues. Certainly, if I want to detect corruption, I use a sha256 of the entire file. If that fails, I don't waste time trying to find the bad bit. To add to that, if you need parity to recover from errors, you need to calculate how much based on your storage medium durability and projected life span. It's not the file format's concern. The xz crc should be irrelevant.
- stefco_ 8y agoWhile that's true for most use cases, I think the author's point is that an archival compression format should be as forgiving as possible to the person recovering data because they are not necessarily the person who stored it. There will certainly be plenty of data in the future that was haphazardly stored but which needs to be recovered, possibly centuries after it was originally created, when no other copies may exist. So we should try to be nice to future archivists/librarians by making our data formats as robust as possible (in addition to our storage media, which is what you are correctly implying we should also worry about).
- pmoriarty 8y ago"If I want to recover a corrupted file, I find another copy." So you've archived two or more copies of each file? That means you're use at least twice as much space (and if you're keeping the original as well, more than twice). For the likely corruption of the occasional single bit flip here and there, you could do a lot better by using something like par2 and/or dvdisaster (depending on what media you're archiving to).
- outworlder 8y agoYes. If your data is not in three different places it might as well not exist.
- jlgaddis 8y ago> So you've archived two or more copies of each file You haven't? It took me just one minor "data loss incident" ~20 years ago to very quickly convince me to become a lifetime member of the "backup all the things to a few different locations" club. > That means you're use at least twice as much space (and if you're keeping the original as well, more than twice). "Storage is cheap."
- carussell 8y ago(2016) Previously discussed here on HN back then: https://news.ycombinator.com/item?id=12768425 https://news.ycombinator.com/item?id=12768425 The author has made some minor revisions since then. Here are the main differences to the page compared to when it was first discussed here: http://web.cvs.savannah.nongnu.org/viewvc/lzip/lzip/xz_inadequate.html?r1=1.3&r2=1.4 http://web.cvs.savannah.nongnu.org/viewvc/lzip/lzip/xz_inade... And here's the full page history: http://web.cvs.savannah.nongnu.org/viewvc/lzip/lzip/xz_inadequate.html http://web.cvs.savannah.nongnu.org/viewvc/lzip/lzip/xz_inade...
- eesmith 8y agoFWIW, PNG also "fails to protect the length of variable size fields". That is, it's possible to construct PNGs such that a 1-bit corruption gives an entirely different, and still valid, image. When I last looked into this issue, it seemed that erasure codes, like with Parchive/par/par2, was the way to go. (As others have mentioned here.) I haven't tried it out as I haven't needed that level of robustness.
- nailer 8y agoTo read the article: document.body.style['max-width'] = '550px'; document.body.style.margin = '0 auto'
- fenwick67 8y agoor just resize your browser window
- jwilliams 8y agoI sent a reasonable amount of data to Cloud Storage. It varies a lot. Usually ~10GB/day, but it can be up to 1TB/day regularly. xz can be amazing. It can also bite you. I've had payloads that compress to 0.16 with gzip then compress to 0.016 with xz. Hurray! Then I've had payloads where xz compression is par, or worse. However, with "best or extreme" compression, xz can peg your CPU for much longer. gzip and bzip2 will take minutes and xz -9 is taking hours at 100% CPU. As annoying as that is, getting an order of magnitude better in many circumstances is hard to give up. My compromise is "xz -1". It usually delivers pretty good results, in reasonable time, with manageable CPU/Memory usage. FYI. The datasets are largely text-ish. Usually in 250MB-1GB chunks. So talking JSON data, webpages, and the like.
- londons_explore 8y agoIf you get compression ratios that good, you should consider if your application might be doing something stupid like storing the same data thousands of times inside it's data file. If you store enough of the same type of data, invest in redesigning the application. There's a reason we all use jpegs over zipped bitmaps...
- UK-Al05 8y agoIt sounds like his application scraping data of some kind rather than say generating it.
- rspeer 8y agoHTML is pretty repetitive, but if you want to archive HTML data, you don't get to redefine what HTML is. Compression is useful.
- chatmasta 8y agoThis is what the WARC [0] file format (and/or gzip) is for. [0] https://en.m.wikipedia.org/wiki/Web_ARChive https://en.m.wikipedia.org/wiki/Web_ARChive
- pmoriarty 8y agoWhen I use xz for archival purposes I always use par2[1] to provide redundancy and recoverability in case of errors. When I burn data (including xz archives) on to DVD for archival storage, I use dvdisaster[2] for the same purpose. I've tested both by damaging archives and scratching DVDs, and these tools work great for recovery. The amount of redundancy (with a tradeoff for space) is also tuneable for both. [1] - https://github.com/Parchive/par2cmdline https://github.com/Parchive/par2cmdline [2] - http://dvdisaster.net/ http://dvdisaster.net/
- londons_explore 8y agoThe purpose of a compression format is not to provide error recovery or integrity verification. The author seems to think the xz container file format should do that. When you remove this requirement, nearly all his arguments become moot.
- zzzcpan 8y ago> The purpose of a compression format is not to provide error recovery or integrity verification. On the contrary. People archive files to save space, exchange files with each other over unreliable networks able to corrupt data, store them in corrupted ram and corrupted disks, even if just temporary. Compression formats are there to help with that, this is their main purpose. This is why fast and proper checksumming is expected, but not cryptographic, like sha256, that adds nothing to this goal but overhead.
- qwerty456127 8y agoWhy do people use xz anyway? As for me I just use tar.gz when I need to backup a piece of a Linux file system into an universally-compatible archive, zip when I need to send some files to a non-geek and 7z to backup a directory of plain data files for myself. And I dream of the world to just switch to 7z altogether but it is hardly possible as nobody seems interested in adding tar-like unix-specific metadata support to it.
- LinuxBender 8y agoxz has substantially better compression than gz or bz2, especially if using the flags -9e. You can use all your cores with -T0 or set how many cores to use. I find it to be on par with 7-zip. Perhaps folks are trying to stick with packages that are in their base repo. p7zip is usually outside of the standard base repos.
- yason 8y agoSubstantially is a relative term. There are niche cases but how many people really care, or need to care, about the last bytes that can be compressed? Packing a bunch of files together as .tgz is a quite universal format and compresses most of the redundancy out. It has some pathological cases but those are rare, and for general files it's still in the same ballpark with other compressors. I remember using .tbz2 in the turn of the millennium because at the time download/upload times did matter and in some cases it was actually faster to compress with bzip2 and then send over less data. But DSL broadband pretty much made it not matter any longer: transfers were fast enough that I don't think I've specifically downloaded or specifically created a .tbz2 archive for years. Good old .tgz is more than enough. Files are usually copied in seconds instead of minutes, and really big files still take hours and hours. None of the compressors really turn a 15-minute download into a 5-minute download consistently. And the download is likely to be fast enough anyway. Disk space is cheap enough that you haven't needed the best compression methods for ages in order to stuff as much data on portable or backup media. Ditto for p7zip. It has more features and compresses faster and better but for all practical purposes zip is just as good. Eventhough it's slower it won't take more than a breeze to create and transfer, and it unzips virtually everywhere.
- leni536 8y agoI fail to see why integrity checking is the file format's responsibility. Is this historical? Like when you just dd a tar file directly onto a tape and there is no filesystem? Anyway seems like it should be handled by the filesystem and network layers. I can understand the concerns about versioning and fragmented extension implementations though.
- JdeBP 8y ago> you just dd a tar file directly onto a tape Actually, one uses the tape archive utility, tar, to write directly to the tape. (-:
- ebullientocelot 8y agoThe [Koopman] cited throughout is my boss, Phil! At any rate I'm sadly not surprised and a little appalled that xz doesn't store the version of the tool that did the compression..
- orbitur 8y agoRelated: where can I find a thorough step-by-step method for maintaining the integrity of family photos/videos in backups on either Windows or macOS?
- LinuxBender 8y agoPerhaps renice your job so that others don't complain about their noisy neighbor. renice 19 -p $$ > /dev/null 2>&1 then ... Use tar + xz to save extra metadata about the file(s), even if it is only 1 file. tar cf - ~/test_files/* | xz -9ec -T0 > ./test.tar.xz If that (or the extra options in tar for xattrs) is not enough, then create a checksum manifest, always sorted. sha256sum ~/test_files/* | sort -n > ~/test_files/.sha256 Then use the above command to compress it all into a .tar file that now contains your checksum manifest.
- moltensyntax 8y agoThis article again? In my opinion, this article is biased. The subtext here is that the author is claiming that his "lzip" format is superior. But xz was not chosen "blindly" as the article claims. To me, most of the claims are arguable. To say 3 levels of headers is "unsafe complexity"... I don't agree. Indirection is fundamental to design. To say padding is "useless"... I don't understand why padding and byte-alignment that is given so much vitriol. Look at how much padding the tar format has. And tar is a good example of how "useless padding" was used to extend the format to support larger files. So this supposed "flaw" has been in tar for dozens of years, with no disastrous effects at all. The xz decision was not made "blindly". There was thought behind the decision. And it's pure FUD to say "Xz implementations may choose what subset of the format they support. They may even choose to not support integrity checking at all. Safe interoperability among xz implementations is not guaranteed". You could say this about any software - "oh no, someone might make a bad implementation!" Format fragmentation is essentially a social problem more than a technical problem. I'll leave it at this for now, but there's more I could write.
- pmoriarty 8y ago"Look at how much padding the tar format has. And tar is a good example of how "useless padding" was used to extend the format to support larger files. So this supposed "flaw" has been in tar for dozens of years, with no disastrous effects at all." Just because it's in tar doesn't mean that the design is flawless. tar was created a long time ago, when a lot of things we are concerned with now weren't even thought of. Deterministic, bit-reproduceable archives are one thing that tar has recently struggled with[1], because the archive format was not originaly designed with that in mind. With more foresight and a better archive format, this need not have been an issue at all. [1] - https://lists.gnu.org/archive/html/help-tar/2015-05/msg00005.html https://lists.gnu.org/archive/html/help-tar/2015-05/msg00005...
- rootbear 8y agoThe name tar comes from Tape ARchive. Lots of padding makes sense when you know that tar was originally used to write files to magnetic tape, which is highly block oriented. The use of tar today as a bundling and distribution format is something of a misapplication, as it lacks features one might want of such a program.
- ryao 8y agoRequiring userland software to worry about bitrot is a great way to ensure that it is not done well. It is better to let the filesystem worry about it by using a file system that can deal with it. This article is likely more relevant to tape archives than anything most people use today.
- AndyKelley 8y agoI did some compression tests of the CI build of master branch of zig: 34M zig-linux-x86_64-0.2.0.cc35f085.tar.gz 33M zig-linux-x86_64-0.2.0.cc35f085.tar.zst 30M zig-linux-x86_64-0.2.0.cc35f085.tar.bz2 24M zig-linux-x86_64-0.2.0.cc35f085.tar.lz 23M zig-linux-x86_64-0.2.0.cc35f085.tar.xz With maximum compression (the -9 switch), lzip wins but takes longer than xz: 23725264 zig-linux-x86_64-0.2.0.cc35f085.tar.xz 63.05 seconds 23627771 zig-linux-x86_64-0.2.0.cc35f085.tar.lz 83.42 seconds
- vortico 8y agoWhat is the probability that a given byte will be corrupted on a hard disk in one year? What is the probability of a complete HD failure in a year?
- sirsuki 8y agoSo what wrong with plain and simple tar c foo | gzip > foo.tar.gz or tar c foo | bzip2 > foo.tar.bz2 Been using these for over 20 years now. Why is is so important to change things especially as this article points out for the worse?!
- dchest 8y agoBetter (smaller and/or faster) compression.
- freedomben 8y agoThis is purely anecdotal and could easily be PEBKAC, but I created a bunch of xz backups years ago and had to access them a couple of years later after a disc died. To my panicked surprise, when trying to unpack them, I was informed that something was wrong (sorry at this point I don't remember what it was). I never did get it working. From that point on I went back to gzip and have not had a problem since. Yes xz packs efficiently, but a tight archive that doesn't inflate is worse than worthless to me.
- comex 8y agoLast time this came up on HN, I did some research, and discovered that lzip was quite non-robust in the face of data corruption: a single bit flip in the right place in an lzip archive could cause the decompressor to silently truncate the decompressed data, without reporting an error. Not only that, this vulnerability was a direct consequence of one of the features used to claim superiority to XZ: namely, the ability to append arbitrary “trailing data” to an lzip archive without invalidating it. Like some other compressed formats, an lzip file is just a series of compressed blocks concatenated together, each block starting with a magic number and containing a certain amount of compressed data. There’s no overall file header, nor any marker that a particular block is the last one. This structure has the advantage that you can simply concatenate two lzip files, and the result is a valid lzip file that decompresses to the concatenation of what the inputs decompress to. Thus, when the decompressor has finished reading a block and sees there’s more input data left in the file, there are two possibilities for what that data could contain. It could be another lzip block corresponding to additional compressed data. Or it could be any other random binary data, if the user is taking advantage of the “trailing data” feature, in which case the rest of the file should be silently ignored. How do you tell the difference? Simply enough, by checking if the data starts with the 4-byte lzip magic number. If the magic number itself is corrupted in any way? Then the entire rest of the file is treated as “trailing data” and ignored. I hope the user notices their data is missing before they delete the compressed original… It might be possible to identify an lzip block that has its magic number corrupted, e.g. by checking whether the trailing CRC is valid. However, at least at the time I discovered this, lzip’s decompressor made no attempt to do so. It’s possible the behavior has improved in later releases; I haven’t checked. But at least at the time this article was written: pot, meet kettle.
- kazinator 8y agoIf the claims in the article are true who cares if the competing thing that the author is working on is also shit (but good to know that too).
- lopmotr 8y agoIt's that an implementation problem? I would expect a decompressor to warn that there's unidentified trailing data and perhaps dump it out as-is. After all, even if you did put it there on purpose, surely you still want it, not to have it discarded.
- microcolonel 8y agoGiven that there is basically one standard implementation, and virtually nobody has ever had an issue with compatibility with a given file, I don't see how it is "inadequate". Sure, if it's inadequate now, it'll be inadequate if you read it in a decade, but not in any way which would prevent you from reading it. If your storage fails, maybe you'll have a problem, but you'd have a problem anyway. Sometimes I feel like genuine technical concerns are buried by the authors being jerks and blowing things way out of proportion. I, for one, tend to lose interest when I hear hyperbolic mudslinging.
- loeg 8y agoUse par2 to generate FEC for your archives and move on with your life.
- Annatar 8y agoSo long as xz(1) gets insane amounts of compression and there is no compressor which compresses better, people are going to keep preferring it.