7 ms·
With all due respect, I find it hard to believe the author stumbled upon a trivial method of improving tarballing performance by several orders of magnitude tha
by xcdzvyn 3y ago
With all due respect, I find it hard to believe the author stumbled upon a trivial method of improving tarballing performance by several orders of magnitude that nobody else had considered before.
If I understand correctly, they're suggesting Pack, which both archives and compresses, is 30x faster than creating a plain tar archive. That just sounds like you used multithreading and tar didn't.
Either way, it'd be nice to see [a version of] Pack support plain archival, rather than being forced to tack on Zstd.
- TylerE 3y agoThat’s more because plain tar is actually a really dumb way of handling files that aren’t going to tape. Being better than that is not a hard bar.
- cogman10 3y agoThe tar file format is REALLY bad. It's pretty much impossible to thread because it's just doing metadata then content and repeatably concatenating. IE /foo.txt 21 This is the foo file /bar.txt 21 This is the bar file That makes it super hard to deal with as you essentially need to navigate the entire tar file before you can list the directories in a tar file. To add a file you have to wait for the previous file to be added. Using something like sqlite solves this particular problem because you can have a table with file names and a table with file contents that can both be inserted into in parallel (though that will mean the contents aren't guaranteed to be contiguous.) Since SQLite is just a btree it's easy (well, known) how to concurrently modify the contents of the tree.
- TylerE 3y agoOr just what zip and every other format does an skits put all the metadata at the beginning - enough to list all files, and extract any single one efficiently
- nullindividual 3y agoTapes don't (? certainly didn't) operate this way. You need to read the entire tape to list the contents. Since tar is a Tape ARchive, the way tar operates makes sense (as it was designed for both output to file and device, i.e. tape).
- monocasa 3y agoTapes currently don't really operate like tar anymore either. Filesystems like LTFS stick the metadata all in one blob somewhere.
- nullindividual 3y agoIt's been a long time since I've operated tape, so good to know things have changed for the better.
- tredre3 3y agoThat point is always raised on every criticism of tar (that it's good at tape). Yes! It is! But it's awful at archive files, which is what it's used for nowadays and what's being discussed right now. Over the past 50 years some people did try to improve tar. People did develop ways to append a file table at the end of an archive file. Maintaining compatibility with tapes, all tar utilities, and piping. Similarly, driven people did extend (pk)zip to cover all the unix-y needs. In fact the current zip utility still supports permissions and symlinks to this day. But despite those better methods, people keep pushing og tar. Because it's good at tape archival. Sigh.
- monocasa 3y agozip interestingly sticks the metadata at the end. That lets you add files to a zip without touching what's already been zipped. Just new metadata at the end. Modern tape archives like LTFS do the same thing as well.
- t43562 3y agoThat sounds like you need to have fetched the whole zip before you can unzip it - which is not what one wants when making "virtual tarfiles" which only exist in a pipe. (i.e. you're packing files in at one end of the pipe and unpacking them at the other)
- sargun 3y agoFunnily enough, tar is like 3 different formats (PaX, tar, ustar). One of the annoying parts of the tar format is that even though you scan all the metadata upfront, you have to keep the directory metadata in RAM until the end and have to wait to apply it at the end.
- darby_eight 3y agoEh, it's not that hard to imagine given how rare it is to zip 81k files of around 1kb each.
- iscoelho 3y agoNot that rare at all. Take a full disk zip/tar of any Linux/Windows filesystem and you'll encounter a lot of small files.
- darby_eight 3y agoOk? How are you comparing these systems to the benchmark so they might be considered relevant? Compressing "Lots of small files" describes an infinite variety of workloads. To achieve anything close to the benchmark you'd need to specifically only compress only small files in a single directory of an average small size. And even the contents of those files would have large implications as to expected performance....
- iscoelho 3y agoMy comment is not making any claims about that. It's just a correction that filesystems with "81k 1KB files" are indeed common.
- darby_eight 3y agoIf that were true, surely it would make sense to demonstrate this directly rather than with a contrived benchmark? The issue is not the preponderance of small files but rather the distribution of data shapes.
- viraptor 3y agoThat's basically any large source repo.
- fbdab103 3y agoZipping up a project directory even without git can be a big file collection. Python virtual environment or node_modules, can quickly get into thousands of small files.
- Hello71 3y agoAlso, 4.7 seconds to read 1345 MB in 81k files is suspiciously slow. On my six-year-old low/mid-range Intel 660p with Linux 6.8, tar -c /usr/lib >/dev/null with 2.4 GiB in 49k files takes about 1.25s cold and 0.32s warm. Of course, the sales pitch has no explanation of which hardware, software, parameters, or test procedures were used. I reckon tar was tested with cold cache and pack with warm cache, and both are basically benchmarking I/O speed.
- lilyball 3y agoThe footnotes at the bottom says > Development machine with a two-year-old CPU and NVMe disk, using Windows with the NTFS file system. The differences are even greater on Linux using ext4. Value holds on an old HDD and one-core CPU. > All corresponding official programs were used in an out-of-the-box configuration at the time of writing in a warm state.
- fbdab103 3y agoHDD for testing is a pretty big caveat for modern tooling benchmarks. Maybe everything holds the same if done on a SSD, but that feels like a pretty big assumption given the wildly different performance characteristics between the two.
- Hello71 3y agoMy apologies, the text color is barely legible on my machine. Those details are still minimal though; what versions of software? How much RAM is installed? Why is 7-Zip set to maximum compression but zstd is not? Why is tar.zst not included for a fair comparison of the Pack-specific (SQLite) improvements on top of from the standard solution?
- OttoCoddo 3y agoUsing 32GB of RAM, but it is far more than they need. 7-Zip was used as others, just gave it a folder to compress. No configuration. As requested, here are some numbers on tar.zst of Linux source code (the test subject in the note): tar.zst: 196 MB, 5420 ms (using out-of-the box config and -T0 to let it use all the cores. Without it, it would be, 7570 ms) Pack: 194 MB, 1300 ms Slightly smaller size, and more than 4X faster. (Again, it is on my machine; you need to try it for yourself.) Honestly, ZSTD is great. Tar is slowing it down (because of its old design and being one thread). And it is done in two steps: first creating tar and then compression. Pack does all the steps (read, check, compress, and write) together, and this weaving helped achieve this speed and random access.
- paulddraper 3y agoIt's like 3x not 30x but yes same skepticism
- jrockway 3y agogzip is really, really, really slow, so it's pretty easy to make a thing that uses gzip fast by switching to Zstandard.
- OttoCoddo 3y agoIt was hard to believe for me, too. And I didn't stumble upon it; I looked for it closely, and that was a point in the note. People did not look properly for nearly three decades. Many things have changed, but we computer people are still using the same tools. I am not saying old is not good; the current solutions are great, but what are we, if we don't look for the better? Yes it is that much faster, and a good part of it is because of the multi-thread design, but as a reminder, WinRAR or 7-Zip are too multi-thread, and you can see the difference. To satisfy your doubt, I suggest running Pack for yourself. I am looking for more data on its behaviour on different machines and data. Can I ask why do you need a version without ZSTD? If you are thinking that compression slows it down, I should say no. Pack is the first of its kinds that "Store" is slowing it down. Because its compression is smart, it will skip any non-compressible content. On the same machine and the same Linux source code test: Pack: 194 MB, 1.3 s Pack (With no Press): 1.25 GB, 1.8 s
- out_of_protocol 3y agoPure zstd (or .tar.zstd) vs pack vs patched 7z+zstd would be more interesting, how much overhead introduced by pack format itself - in size and speed
- OttoCoddo 3y agoI answered this question here: https://news.ycombinator.com/item?id=39801083 https://news.ycombinator.com/item?id=39801083 If that is not enough, let me know.
- out_of_protocol 3y agotar.zst vs pack is looking great, thanks! Also there is https://github.com/mcmilk/7-Zip-zstd https://github.com/mcmilk/7-Zip-zstd .pack vs zst-7z with the same compression settings would b interesting. That will be pure container overhead
- xcdzvyn 3y agoMy concern with Pack obliging me to compress is that compression becomes less pluggable; I'd much rather my archive format be agnostic of compression, as with tar, so that I can trivially move to a better compression format when one inevitably comes to be.