3 ms·
I'm confused by how this affects range requests. Without compression, those can be easily satisfied by reading the relevant part of the cached complete file. Bu
by CodesInChaos 1mo ago
I'm confused by how this affects range requests. Without compression, those can be easily satisfied by reading the relevant part of the cached complete file. But how are they handled now? The article claims "range requests remain unchanged", but I don't see how that's possible if the cache no longer stores the uncompressed data.
- pkulak 1mo agoI assume the entire resource needs to be decompressed first, then indexed into, served, and discarded. Well, actually, you could just decompress up to the end of the range.
- CodesInChaos 1mo agoWhich would have terrible performance for range requests starting late in a large file. For files that are frequently accessed that way, this could be prohibitive. You could split the file into independently compressed blocks as well. But that'd reduce compression rate and require adding some kind of index for seeking. Or they have an upper size limit for the file size they compress, since large files are rarely compressible text. In any case it is something that needs the be handled before going live with a compressed cache. But the article sounds like they simply didn't implement compressed caching for those cases, which makes no sense.
- mgerdts 1mo agoFor the cost of a small amount of metadata the offsets of every MiB or so could be stored. I did something like this with pigz as I was implementing multithreaded compressed and encrypted kernel zone suspend and resume for Solaris.
- genxy 1mo agoNot with zstd, you could still support range requests. https://en.wikipedia.org/wiki/Zstd https://en.wikipedia.org/wiki/Zstd this whole subthread should take 10 minutes and glance over the spec and the capabilities. It would end a lot of wasted premature pontificating.
- ncruces 1mo agoThere's nothing about random access at that link. There's this, but it doesn't seem to be getting much traction: https://github.com/facebook/zstd/tree/dev/contrib/seekable_format https://github.com/facebook/zstd/tree/dev/contrib/seekable_f...
- genxy 1mo agoAlso external dictionaries https://nigeltao.github.io/blog/2022/zstandard-part-7-dictionaries.html https://nigeltao.github.io/blog/2022/zstandard-part-7-dictio...
- kccqzy 1mo agoActually zstd internally splits data into frames, and frames can indicate the decompressed data size. So if we control the compressor we can make it so that all frames have the size information; it isn’t exactly seekable but at least it will not need to decompress the resource. https://python-zstandard.readthedocs.io/en/latest/concepts.html https://python-zstandard.readthedocs.io/en/latest/concepts.h... Given how fast zstd can decompress, this may or may not actually be a win: the time spent waiting for I/O might be so large that the decompression can fit within the wait time.
- gopalv 1mo ago> I don't see how that's possible if the cache no longer stores the uncompressed data. Zstd has a seekable format for frames, similar to pigz --independent works. [1] - https://github.com/facebook/zstd/blob/dev/contrib/seekable_format/README.md https://github.com/facebook/zstd/blob/dev/contrib/seekable_f...
- tehlike 1mo agoYes. This.
- NostraDavid 1mo agoSpeaking of pigz, I've run into pigzpp[1], or rather its paper: "pigzpp: Fast, Parallel, Portable Compression for the Whole Stack"[2]. Turns out we can squeeze quite a bit more compression performance out of DEFLATE - ~10x in certain instances, 2x as a base minimum (read the paper for details). [1]: https://github.com/thammegowda/pigzpp https://github.com/thammegowda/pigzpp [2]: https://arxiv.org/abs/2608.24153 https://arxiv.org/abs/2608.24153
- nijave 1mo agoIdk but btrfs and zfs manage to pull it off Seekable OCI (SOCI) uses an index so I imagine that's an option (real byte range a-b maps to compressed range x-y). Presumably you'd still need to read the header and some additional pieces
- a_t48 1mo agoIt uses some form of keyframing. Not entirely free.
- butvacuum 1mo agoafaik, ZFS will read an entire Record at a time- and that's the same granularity as its compression.
- mgerdts 1mo agoZFS compresses recordsize or volblocksize chunks down to some whole number of disk blocks, as determined by 1 >> ashift. In practice, this typically means that each 128k chunk gets compressed to some number of sequential 512 or 4096 byte blocks. These compressed blocks are referenced by block pointers that contain flags indicating compression and what type.
- butvacuum 1mo agothey said they didn't change the behavior for range requests. So, it'll still be the basic no compression. (eg, server side decompression)