5 ms·
Sure, but now you've added an extra layer of indirection which can have a significant impact on performance
by jeffffff 4y ago
Sure, but now you've added an extra layer of indirection which can have a significant impact on performance
- klauspost 4y agoIt doesn't really have to impact performance. The index is generated easily as a side-effect of compression. And the index is only needed if you need to seek. I implemented this as part of the MinIO server. See "Seeking Compressed Files" here: https://blog.min.io/transparent-data-compression/ https://blog.min.io/transparent-data-compression/ We choose a compressor without literal compression for a faster baseline, but the concept remains the same.
- jeffffff 4y agoBut if you do need to seek, which is really common in data warehouse workloads for example, unless you keep the index in ram you have to do an extra IO on every seek to read the index
- blibble 4y agothere's always going to be some metadata for the file that needs to be looked up before you can start seeking (ACLs, sector/extent/cluster location, etc) the index goes in there, no extra seek needed
- jeffffff 4y agoyeah that isn't free either, it adds significant bloat to your metadata. with most enterprise customers encrypting and/or compressing data before putting it into s3, it doesn't seem like there would be much benefit. s3 really isn't the right layer to implement compression. filesystems aren't either. it's better to leave it up to the application.
- blibble 4y ago> yeah that isn't free either, it adds significant bloat to your metadata yeah, 4 bytes for every megabyte > s3 really isn't the right layer to implement compression. filesystems aren't either. it's better to leave it up to the application. yeah, I'm sure you're right and Amazon have absolutely no idea what they're doing and like to spend unnecessary CPU cycles doing pointless work and add "significant bloat" to their metadata ... or, you're wrong (like in every previous comment in this chain)
- jeffffff 4y agohttps://www.reddit.com/r/programming/comments/wtd61q/aws_switch_from_gzip_to_zstd_about_30_reduction/il4uu67/ https://www.reddit.com/r/programming/comments/wtd61q/aws_swi... this tweet is not talking about compressing customer data in s3, i seriously doubt that aws compresses customer data in s3 for all the reasons i've already listed. i am right and amazon does know what they're doing, which is why they don't compress customer data in s3. 4 bytes per megabyte becomes significant at scale when you have to keep it in ram, which you have to do if you want to avoid the extra IO.
- blibble 4y agoah yes, "authoritative" comments from random reddit accounts and you don't understand the algorithm if you think you need to keep the index in RAM, because you don't
- jeffffff 4y agoif it's not in ram, you have to do an extra IO to look it up. i don't think you understand how precious metadata space is in a large scale storage system. if you pollute the metadata cache with useless junk like this, you can't cache as many things, your hit rate goes down, and you have to do more IO operations to service each request on average. name one popular distributed file system or object store that compresses everything by default like you are claiming. you won't be able to, because none of them do it, because it's better to leave it to the application.