4 ms·
Had not considered that. Reading about squashfs I see that it's readonly. Do you know if that's fine when mounting the directory for an LXC container? Also I wo
by ctecte 5y ago
Had not considered that. Reading about squashfs I see that it's readonly. Do you know if that's fine when mounting the directory for an LXC container? Also I wonder how much performance impact (if any) there might be on spark jobs due to the decompression at runtime.
To be completely honest this all started as a quick hackathon project just to speed up downloading before piping to GNU tar. Only later did I consider also re-implementing the tar extraction in a parallel manner. If considering changing the overall packaging method from tar to something else there's a lot more ways to consider going about this (including squashfs).
I'd say one of the nice things about this is that tar is fairly ubiquitous, not just in container images. For example, we also have tarball build artifacts in Jenkins jobs that could benefit from this tool.
- tarasglek 5y agoYou can cut down the startup time to milliseconds 1) either mount tar directly 2) or use squashfs and mount that...both would be mounted with some sort of rw overlay 3) underneath do some lazy httpfs type filesystem, so filesystem can be mounted while it's downloading 4) parallelize the underlying download ala aria2c 5) provide metadata to aria2c-alike downloader with file boundaries, so extract further latency/bandwidth savings
- ctecte 5y ago1) Are there any ways to do this in a performant manner? Afaik, since the tarball is still sequential with no way to jump around, interacting with this filesystem would be fairly slow right? 2) Agreed, still need to look more into this, although it's more involved and entails changing more of our pipeline for packaging and distributing these images. 3) I need to look into how this would handle files that haven't been downloaded yet. From what I know about httpfs filesystems, I'm not sure how much would need to be done to let this block until file needed is downloaded vs the normal behavior of calling out to get the file being requested. 5) Could you go into more detail here? Not sure I understand how file boundaries, etc can help with latency/bandwidth.
- hawski 5y agoAd 3. With squashfs via httpfs it would probably be fast enough. I remember years ago httpfs capable of booting livecd iso and it was performant enough. I'm not current with this knowledge, but is there a httpfs, that would download increasingly cache the file while doing range requests as it gets read requests for certain parts of the file?
- tarasglek 5y ago3) it just blocks and prioritizes that part of download 5) if you can chuck the file so chunks that get downloaded are whole files..less likely to do multiple requests for a single file 6)I forgot to do the best part of the optimization..do feedback-guided-optimization. You can then repack the squashfs/tar file to have all the frequently-accessed files together
- ctecte 5y agoAhh got it, this makes a lot of sense, maybe something to try next hackathon :)
- ffk 5y agoSomething to consider, if you are IO constrained, compression may speed up reads because you shift some of the cost of IO to the CPU. Ultimately, you'll need to measure this to know for sure, and those results will likely only be valid on a given hardware configuration. OverlayFS also has a "copy_up" function, where the file is copied at the initial write. Once the copy is done, I'd expect write access to be fast. Again, you'll need to measure this. The setup could probably look like: container read/write -> OverlayFS([mutable fs as overlay] -> [squashfs layer as underlay] -> [squashfs layer as underlay])