10 ms·
There's a balance between compression ratio and CPU utilisation. Zstd is several times faster at compressing than gzip to produce a file of similar size (it's g
by IMSAI8080 4y ago
There's a balance between compression ratio and CPU utilisation. Zstd is several times faster at compressing than gzip to produce a file of similar size (it's good, give it a try if you haven't already). I guess he means by using zstd they were able to crank up the compression ratio and maintain the same CPU usage maybe?
I don't think AWS routinely compress customer data that I've noticed. I guess he must mean for their internal products that use S3 perhaps?
- notimetorelax 4y agoI doubt there’s a single byte stored to disk that is not compressed and encrypted at AWS. It’s transparent to the customer.
- deleted 4y ago[deleted]
- smueller1234 4y agoIt's not quite that simple. If you have customer data that's already encrypted, then compression won't do much because it looks random. But of course by the time you get to your infrastructure layers, that'll be the case (or you really messed up your security story!). Which means you'd have to compress right at the edge. They might be doing that (which would basically mean it's the customer compressing it before they encrypt it with their keys because AWS has no business seeing the clear text), but then you get to compress each item separately, which might not be very effective for small values. tl;dr:There's a real efficiency/security/insider risk trade-off here. Edit: I should disclose that I work for a competitor. Don't intend any astroturfing.
- notimetorelax 4y agoI agree with you, there could be scenarios where customers supply their own keys and compress the data on their own. My original statement is still true though, the data at rest ends up being compressed and encrypted. That said, of course, customers can upload encrypted blobs of uncompressed data. But I’d call it an exception that proves the rule. Here service simplicity should win and those blobs may end up recompressed.
- alexchamberlain 4y ago+1 if you are storing objects uncompressed, I'd be amazed if AWS doesn't compress them and charge you for the full space anyway
- sitkack 4y agoIf this true, there is possibly a side channel one could run against object storage to determine if someone else in the content-addressable-store has the same files. Like when it was easy to file share on dropbox by having the correct hashes. A GUID could summon a 1GB file.
- staticassertion 4y agoThat assumes cross-tenant compression.
- eurg 4y agoCompression and content-addressing are two separate things. Content addressing across accounts on private, AWS encrypted S3 buckets would run counter to their claims.
- alexchamberlain 4y agoA couple of comments across the thread have made similar points, but if I were implementing this, the "client metadata" like the incoming sha256 etc would be implemented a layer higher than the actual byte storage, so the byte storage could be compressed without any impact on that sort of thing.
- IMSAI8080 4y agoI don't think it's so clear cut. They have to pay to compress it. If the data the customer stores is short lived it may not be worth it to them. They don't know if the customer already compressed it so they might be wasting their CPU. They also have to pay to decompress it on every access. They allow you to slice an arbitrary byte range out of an object which is technically harder to implement on a compressed file. They charge by the GB and are not exactly super cheap so if the customer wants to store a big fat file of easily compressible zeros then whatever, they got their money. It might make more sense on their "deep archive" product maybe where the customer has to commit to a minimum storage retention and also pay a retrieval charge which scales with the amount of data recovered (hence paying for the CPU to decompress).
- mekster 4y agoWhy do they have to either compress it all or not. They must be smart like, have the files split in pieces (just like some network file systems/backups do) and if those blocks are untouched for a while, compress what's compressible and leave them as is during active uses.
- pythux 4y ago> They charge by the GB and are not exactly super cheap so if the customer wants to store a big fat file of easily compressible zeros then whatever, they got their money. And what if they can charge you the uncompressed size and only actually store much smaller compressed files behind the scene. That seems like having your cake and eat it too.
- notimetorelax 4y agoThis was true a few years back, nowadays it’s cheaper and faster to compress the data at rest as the bottleneck is frequently IO and storage space. Both, in terms of capacity and cost.
- microtonal 4y agoThis was true a few years back Only temporarily with SSDs. With spinning rust, it also often paid off to compress data. We'd store large treebanks compressed, because decompression was much faster than disk reads.
- deleted 4y ago[deleted]
- HiJon89 4y agoHow would that work for something like S3 range requests? Rather than reading an entire object sequentially (which would work fine with transparent compression) you can also ask to read an arbitrary byte range (give me bytes 1,000,000,000-1,000,001,000 from the original file). I guess you could maybe store the compressed file in chunks with metadata about the original byte range inside each chunk.
- rcxdude 4y agoGenerally with filesystem-level compression you don't compress an entire multi-GB file: you compress segments of maybe a few 100k. This gives you a very slightly worse compression ratio but allows random seeks to still be efficient.
- klauspost 4y agoFor MinIO (an S3 compatible server), we add an index for each part, which contains uncompressed -> compressed offset pairs. Since we already used a Snappy-derived method, each 1MB block is stored without backreferences. With this we only have to decode at most 1MB-1 extra bytes to respond with a specific range offset.
- uluyol 4y agoI'm sure it's encrypted, but I doubt that they compress everything. Images and video tend not too compress well since they've typically already been aggressively compressed with specialized algorithms. It would just be throwing CPU cycles away.
- usefulcat 4y agoIt’s also faster to decompress. So it would likely reduce net CPU use for read-mostly resources.
- zxcvbn4038 4y agoThey do compress for most log delivery types like load balancers, cdn, cloudtrail, etc. and it makes a huge difference over raw logs - compression ratios are in the 90s. One AWS specific trick is individual log files below some size threshold are passed through raw, so you end up with a mix of compressed and uncompressed files in S3, and the only reliable way to distinguish between them is to receive them and look for a zlib header - you can’t depend on the file name or any other metadata to tell you ahead of time if an individual file is compressed or not. (I think Cloudtrail does set metadata correctly for uncompressed files, but other types do not, best to spend the time developing a abstraction that deals with both)
- whoknew1122 4y agoAmazon is AWS's biggest customer, but I've never been told to compress things with zstd while working at AWS. And services still compress files as gzip before storing data in S3. Maybe data in S3 is compressed? I don't know the internals of S3.
- buchanmilne 4y ago> Amazon is AWS's biggest customer, but I've never been told to compress things with zstd while working at AWS. You may have switched transparently to using it (e.g. for your application log files), via internal Amazon tooling (that many AWS services use). However, this tweet wasn't about AWS customer's data, but AWS's own data. > Maybe data in S3 is compressed? I don't know the internals of S3. Almost every other AWS service uses S3 for storing something. Think of EBS snapshots, DynamoDB backups, RDS backups. And that's not even the service-internal data.