5 ms·
"... there are some metadata extensions that allow this)." Where to find these extensions? Are they portable between Linux and BSD? The 1998 dict project inc
by textmode 8y ago
"... there are some metadata extensions that allow this)."
Where to find these extensions? Are they portable between Linux and BSD?
The 1998 dict project included a utility called "dictzip" for random access to the contents of gzip compressed files.
Dumb question: Is it possible to create a utility or even a hack that performs "random access" into tar archives?
Example use case: the user only wants to untar a small number of selected files from a large tarball such as a source tree.
The user has tried both the "-T filelist" option and using memory file systems instead of hard disk drives.
- Hello71 8y agoafaik, only pixz does this; they store the index in xz.
- barrkel 8y agoA zip file is a concatenation of gzipped files. A .tar.gz is a gzip stream of concatenated files. Anything that could do random access into the contents of a zip file entry could do similar things with a tarball.
- jahewson 8y agoNot so simple. A bundle of streams is not the same as a stream of bundles.
- sakuronto 8y agoWhat about .gz.tar, in which all the files are gzipped, then tarballed? It seems like it would be a slightly fatter .zip file.
- barrkel 8y agoWith a transparent random access overlay, the difference mostly disappears, reducing to whether the stream needs to be scanned or whether it's indexed, which is itself orthogonal - zip file directory at the end is redundant.
- netheril96 8y agoSo you mean at each "random access", you actually have to scan the whole .tar.gz file to find the location? For large tarballs, that will definitely hinder performance a lot. The difference does not disappear at all.
- barrkel 8y agoApparently you have a comprehension problem.
- netheril96 8y agoThen enlighten me.
- deleted 8y ago[deleted]
- gkfasdfasdf 8y agoZip files have a directory which lists the files in the archive and their offsets, etc. No such feature in a tar archive.
- barrkel 8y agoIf we're talking about something that indexes gzip streams, it's not a leap for it to also index the inner tar.
- Mikhail_Edoshin 8y agoAFAIK a compressor like zip builds a dynamic running table of frequent byte sequences; the resulting archive is written in such a way that when you decompress it, you re-build the table in the process. So if you concatenate files A, B, and C and then compress the result, then by the time the compressor starts compressing the data of C, it will have that table built from A and B. To extract C, you'll need to re-build the same table and thus you'll need first to decompress A and B. In a zip file each entry is compressed individually; this gives random access, but worse compression rate, because the table is not re-used between files.
- pwg 8y ago> Dumb question: Is it possible to create a utility or even a hack that performs "random access" into tar archives? Yes. If the tar file is on a seek-able medium, just read the headers only and build an index to the offsets of the file contents from the header data. Then use the offsets index to seek to just the item of interest and read out only it and nothing more. Now, does such a utility already exist? The answer seems to be yes: https://github.com/devsnd/tarindexer https://github.com/devsnd/tarindexer