5 ms·
There's one thing I don't understand. Each time a new compression algorithm is introduced, it's the Next Big Thing. Why isn't the implementation of the algorith
by d33 6y ago
There's one thing I don't understand. Each time a new compression algorithm is introduced, it's the Next Big Thing. Why isn't the implementation of the algorithm as simple as linking in the related library, assuming they'd all have a similar interface? After all, it seems like what you need is a header and a function that converts a compressed block to a decompressed one and the other way round. Where's the complexity from the API/implementation perspective given that a library is already there?
- georgyo 6y agoFor one thing, kernel models don't link against libraries. So a new implementation is almost always written.
- Someone 6y agoThis incorporates the zstd source code ‘as is’. FTA: Frequently Asked Questions “Q: Why is it so many lines of code, I can't review all of that... A: Most of this code is the ZSTD library, which has not been altered.” One reason this needs testing because this is a file system. If it breaks, it can lose you much more data than the file being worked on. There also may be serious performance degradation on hardware or configurations the developers didn’t look at. There also is some new code added to call the zstd library.
- tobias3 6y agoIn this case it could use the Linux crypto API, which already has ZSTD support and provides compress/decompress functions. But that is exported as GPLv2 so ZFS needs to do its own thing. And idk if the API is sufficient w.r.t. to e.g. workspace management/reuse. W.r.t. to ZSTD: The usual thing is to provide a zlib compatible API, which ZSTD does ( https://github.com/facebook/zstd/tree/dev/zlibWrapper https://github.com/facebook/zstd/tree/dev/zlibWrapper ). But the ZSTD zlib compatibility layer causes lower performance, so it is better to use it directly. Maybe all future compression algorithms can provide a ZSTD compatible API, so we have to do less work in the future?
- m4rtink 6y agoIt often works that way - when binary RPM compression switched from xz to ZSTD in Fedora recently, it was reaaly just about making sure the tooling can now read ZSTD while keepin xz support in place for compatibility (basically msking sure it links to ZSTD and can use it at runtime). Then just switch the Fedora build system to compress binary RPMs with ZSTD when they are built and you are done. :)
- microcolonel 6y ago> Why isn't the implementation of the algorithm as simple as linking in the related library Because filesystems need to store compression metadata and compressor settings differently than individual archive streams do. Because of the way ZFS stores configuration like this, in the previous version of these patches, they had to choose a subset of the available compression levels when adding zstd support. Different compressors and decompressors also have different state sizes for different settings. Allocating, reusing, and discarding buffers for compression/decompression state in a sensible way inside an operating system kernel is not trivial.
- vkaku 6y agoIt's usually the same as linking, except there are tunables in every library, and the hard part is figuring out which one is appropriate for the job. Maybe, this call is to ask people how this library works out in the real world, on the current datasets that people have. Besides, they are talking of testing implementation specific limitations, and may they want to test for that.