4 ms·
You only have to decompress it first if it's compressed (commonly using gzip, which is shown with the .gz suffix). Otherwise, you can randomly access any file
by nrclark 5mo ago
You only have to decompress it first if it's compressed (commonly using gzip, which is shown with the .gz suffix).
Otherwise, you can randomly access any file in a .tar as long as:
- the file is seekable/range-addressible
- you scan through it and build the file index first, either at runtime or in advance.
Uncompressed .tar is a reasonable choice for this application because the tools to read/write tar files are very standard, the file format is simple and well-documented, and it incurs no computational overhead.
- electroly 5mo agoYou've just constructed your own crappy in-memory zip file, here. If you have to build your own custom index, you're no longer using the standard tools. If you find yourself building indices of tar files, and you control the creation, give yourself a break and use a zip file instead. It has the index built in. Compression is not required when packing files into a zip, if you don't want it.
- marginalia_nu 5mo agoYeah it's pretty common to use zip files as purely a container format, with no compression enabled. You can even construct them in such a way it's possible to memory map the contents directly out of the zip file, or read them over network via a small number of range requsts.
- johannes1234321 5mo ago> Uncompressed .tar is a reasonable choice for this application Yes, uncompressed tar (with transfer compression, which is offered in HTTP) is an option for some amount of data. Till the point where it isn't. zip has similar benefits as tar(+transfer compression) but a later point where it fails for such a scenario.
- chungy 5mo agoZip allows you to set compression algorithm on a per-file basis, including no compression.
- QuantumNomad_ 5mo agoYou can achieve the same with tar if you individually compress the files before adding them to the tar ball instead of compressing the tar ball itself. I don’t see how that plus a small index of offsets would be notably more or less work to do from using a zip file.
- chungy 5mo agoZip has a central directory you could just query, instead of having to construct one in-memory by scanning the entire archive. That's significantly less work.
- QuantumNomad_ 5mo agoI mean if they include a pre-made index with it. For example an uncompressed index at byte offset 0 in the tar ball that lists what is inside and their offsets. It would still be comparable amount of work to create software to do that with tar as to use a zip file, if fine grained compression levels etc is being used.
- johannes1234321 5mo agoBut then you are not using tar, you are doing your own file format atop of tar.
- QuantumNomad_ 5mo agoI suppose you are right about that. But it would still be a valid tar file that can be viewed and extracted with normal tools. Kind of similar to how a .docx file can be extracted as zip but still has additional structure to its contents.
- chungy 5mo agoWhat are you really proposing? That a first ".INDEX" entry be made that contains the offsets of all the other members? That could work in a backwards-compatible way (as long as no standard tar utility makes modifications to the archive...), but it's hamfisted. Just use Zip. It's already a well-known format with numerous implementations and already does the job that you want to do.
- kevin_thibedeau 5mo agoRomfs is more capable, simple to support, and doesn't have the overhead of tar's large headers and typical large blocking factors.