4 ms·
Mounting tar archives as a filesystem in WebAssembly
- jghn 6mo agoIsn't "archive" embedded in "tar" already? In other words, is this like saying one went to the "ATM machine"?
- sillysaurusx 6mo agoOnly peripherally relevant, but also see Ratarmount: https://github.com/mxmlnkn/ratarmount https://github.com/mxmlnkn/ratarmount It lets you mount .tar files as a read only filesystem. It’s cool because you basically get random access to the tarball without paying any decompression costs. (It builds an index saying exactly where so-and-so is for every file.)
- jsrcout 5mo agoRatarmount is so cool. Recently I wanted to look at a couple random files in a > 300G compressed tarball. It just wouldn't have been worth doing without it.
- Ecco 6mo agoHow about using a format that has actually been designed to be a compressed read-only filesystem? Something like a SquashFS or cramfs disk image?
- stingraycharles 6mo agoSometimes (read: very often) you can’t choose the format. Obviously if squashfs is available that is a better solution.
- johannes1234321 6mo agoWhen looking at established file formats, I'd start with zip for that usecase over tarballs. zip has compression and ability to access any file. A tarfule you have to uncompress first. SquashFS or cramps or such have less tooling, which makes the usage for generating, inspecting, ... more complex.
- blipvert 6mo agoZip is a piece of cake. I had need to embed noVNC into an app recently in Golang. Serving files via net/http from the zip file is practically a one-liner (then just a Gorilla websocket to take the place of websockify).
- nrclark 6mo agoYou only have to decompress it first if it's compressed (commonly using gzip, which is shown with the .gz suffix). Otherwise, you can randomly access any file in a .tar as long as: - the file is seekable/range-addressible - you scan through it and build the file index first, either at runtime or in advance. Uncompressed .tar is a reasonable choice for this application because the tools to read/write tar files are very standard, the file format is simple and well-documented, and it incurs no computational overhead.
- electroly 6mo agoYou've just constructed your own crappy in-memory zip file, here. If you have to build your own custom index, you're no longer using the standard tools. If you find yourself building indices of tar files, and you control the creation, give yourself a break and use a zip file instead. It has the index built in. Compression is not required when packing files into a zip, if you don't want it.
- phiresky 6mo agoI'm a bit disappointed that this only solves the "find index of file in tar" problem, but not at all the "partially read a tar.gz" file problem. So really you're still reading the whole file into memory, so why not just extract the files properly while you are doing that? Takes the same amount of time (O(n)) and less memory. The gzip-random-access problem one is a lot more difficult because the gzip has internal state. But in any case, solutions exist! Apparently the internal state is only 32kB, so if you save this at 1MB offsets, you can reduce the amount of data you need to decompress for one file access to a constant. https://github.com/mxmlnkn/ratarmount https://github.com/mxmlnkn/ratarmount does this, apparently using https://github.com/pauldmccarthy/indexed_gzip https://github.com/pauldmccarthy/indexed_gzip internally. zlib even has an example of this method in its own source tree: https://github.com/gcc-mirror/gcc/blob/master/zlib/examples/zran.c https://github.com/gcc-mirror/gcc/blob/master/zlib/examples/... All depends on the use case of course. Seems like the author here has a pretty specific one - though I still don't see what the advantage of this is vs extracting in JS and adding all files individually to memfs. "Without any copying" doesn't really make sense because the only difference is copying ONE 1MB tar blob into a Uint8Array vs 1000 1kB file blobs One very valid constraint the author makes is not being able to touch the source file. If you can do that, there's of course a thousand better solutions to all this - like using zip, which compresses each file individually and always has a central index at the end.
- ImJasonH 6mo agoThe first time I'd heard of this was via https://github.com/jonjohnsonjr/dagdotdev/blob/main/internal/explore/README.md#blobs https://github.com/jonjohnsonjr/dagdotdev/blob/main/internal... which powers https://oci.dag.dev https://oci.dag.dev to let you browse OCI images (e.g., https://oci.dag.dev/fs/ubuntu@sha256:b40150c1c2717d324cdb17278c8efdfa4dfcd2ffe083e976f0bcedf31115f081/?mt=application%2Fvnd.oci.image.layer.v1.tar%2Bgzip&size=29732978 https://oci.dag.dev/fs/ubuntu@sha256:b40150c1c2717d324cdb172...)
- a_t48 6mo agoThis is very cool. Worth a submission by itself.
- Lerc 6mo agoI did some similar shenanigans when I did a silly little system on NeoCities https://lerc.neocities.org/ https://lerc.neocities.org/ It uses IndexedDB for the filesystem. Rather Dumbly it is loading the files from a tar archive that is encoded into a PNG because tar files are one of the forbidden file formats.
- haunter 6mo agoNow I want to try how does that work with BTFS which in a similar vein mounts a torrent file or magnet link as a read only directory https://github.com/johang/btfs https://github.com/johang/btfs
- crabique 6mo agoVery cool, I wish there were something similar to this for filesystem images though. Just recently I needed to somehow generate a .tar.gz from a .raw ext4 image and, surprisingly, there's still no better option than actually mounting it and then creating an archive. I managed to "isolate" it a bit with guestfish's tar-out, but still it's pretty slow as it needs to seek around the image (in my case over NBD) to get the actual files.
- tredre3 6mo agoThere are surprisingly few tools to work on file system images in the Linux world, they expect loopback mounting to always be available. There are a few libraries to read ext4 but every time I've tried to use one it missed one feature that my specific image was using (mke2fs changes its defaults every couple years to rely on newer ext4 features). 7-zip can also read ext4 to some degree and, I'm not sure but, they seem to have written a naive parser of their own to do it: https://github.com/mcmilk/7-Zip/blob/master/CPP/7zip/Archive/ExtHandler.cpp https://github.com/mcmilk/7-Zip/blob/master/CPP/7zip/Archive...
- Dwedit 6mo agoTAR archives are good in a few ways, but random access to files is not one of them. You need to iterate over every file before you can create a mapping between filename and its TAR file address. (Meanwhile, sending TAR over Netcat is a valid way to clone a filesystem to another computer, including maintaining the hardlinks and symlinks)
- biglio23 6mo ago[dead]