4 ms·
I don't think it's so clear cut. They have to pay to compress it. If the data the customer stores is short lived it may not be worth it to them. They don't know
by IMSAI8080 4y ago
I don't think it's so clear cut. They have to pay to compress it. If the data the customer stores is short lived it may not be worth it to them. They don't know if the customer already compressed it so they might be wasting their CPU. They also have to pay to decompress it on every access. They allow you to slice an arbitrary byte range out of an object which is technically harder to implement on a compressed file. They charge by the GB and are not exactly super cheap so if the customer wants to store a big fat file of easily compressible zeros then whatever, they got their money.
It might make more sense on their "deep archive" product maybe where the customer has to commit to a minimum storage retention and also pay a retrieval charge which scales with the amount of data recovered (hence paying for the CPU to decompress).
- mekster 4y agoWhy do they have to either compress it all or not. They must be smart like, have the files split in pieces (just like some network file systems/backups do) and if those blocks are untouched for a while, compress what's compressible and leave them as is during active uses.
- pythux 4y ago> They charge by the GB and are not exactly super cheap so if the customer wants to store a big fat file of easily compressible zeros then whatever, they got their money. And what if they can charge you the uncompressed size and only actually store much smaller compressed files behind the scene. That seems like having your cake and eat it too.
- notimetorelax 4y agoThis was true a few years back, nowadays it’s cheaper and faster to compress the data at rest as the bottleneck is frequently IO and storage space. Both, in terms of capacity and cost.
- microtonal 4y agoThis was true a few years back Only temporarily with SSDs. With spinning rust, it also often paid off to compress data. We'd store large treebanks compressed, because decompression was much faster than disk reads.
- deleted 4y ago[deleted]
- LinAGKar 4y agoIt would still produce some CPU overhead, and thus some energy usage.
- IntelMiner 4y agoPresumably it's the tradeoff of CPU overhead versus disk and bandwidth (larger files take longer to copy into memory, which is also energy usage. And more bandwidth to shunt around Amazon's own network)
- SuperQue 4y agoThere's also a latency component. Since CPUs are fast enough to deflate in real-time now, your bottleneck for a read is your storage/network. Reducing the bytes read from storage improves the IO latency.
- deleted 4y ago[deleted]
- erk__ 4y agoThey could be using hardware compression which can be orders of magnitude faster than doing it on the CPU. Hardware compression is sadly not widely available, I think the only consumer product I know with it is the PlayStation 5. The mainframes from IBM have had hardware zlib since Z14 iirc and in my small tests it is very fast compared to the CPU implementation
- sigmoid10 4y agoI think a lot if datacenter SSDs already come with in-drive hardware compression these days, since it not only increases speed but also longevity. So it would actually save money anyways.
- snoopy_telex 4y agoThey do not. It would be difficult to plan correctly if your free disk space is… variable. Example: You have an existing 40 gigabyte file It happened to compress well You delete it and your free disk space goes up by 4 gigabytes. You then write a new 40 gigabyte file that doesn’t compress well Replacing an existing file of the same size just ate an extra 36 gigabytes. How would you plan around that? SSDs should store the bytes given and don’t play fancy games.
- anamexis 4y agoYes they do. https://www.intel.com/content/www/us/en/support/articles/000006354/memory-and-storage.html https://www.intel.com/content/www/us/en/support/articles/000...
- wmf 4y agoNote that these are pre-2017 consumer SSDs. I think SSD compression fell out of favor due to the rise of FDE.
- ChrisLomont 4y ago
- Havoc 4y ago>. They charge by the GB and are not exactly super cheap so if the customer wants to store a big fat file of easily compressible zeros then whatever, they got their money. Maybe they charge at uncompressed rate but store it at compressed? Then they got even more money!
- treffer 4y agoI think it is a clear cut, mostly because I do not think that compression compromises any of those features, all while making the user experience better. For any storage system like this you usually have a few bottlenecks. IO and Network are the obvious ones, followed by tiering (cache, fast io, slow io, ...) and at the very end CPU. Now let's say network is your bottleneck. If you can send the data to the client in a compressed for then you get the compression ratio as additional bandwidth. And the user would get the data quicker! So compression to the network is a clear win. But the common bottleneck is often IO, a high end SSDs with 1M IOPS at 4KB would _theoretically_ serve 4GB/s, a 40GBit link. That's without any redundancy over other overhead. Again compression to the storage layer would decrease the total amount of IOs, thus making sure a customer gets data quicker. Ok, let's say both are not the issue. The fastest compression algorithms compete with memcopy. So if you need just one copy of your data you might have been faster by compressing it. Especially fast compression algorithms (zstd, lz4, snappy, lzo, ...) are worth the CPU cost with virtually no downsides. The problem is finding the right sweet spot that reduces the current bottleneck without creating a CPU bottleneck, but zstd offers the greatest flexibility there, too. Oh for range requests.... Those large objects are likely split anyway, for easier error recovery (imagine 100MB into a 1GB transfer you notice that the file data was corrupted - not good). Once you work on blocks it's easy to do somewhat efficient range requests again.
- dylan604 4y agoHow clear cut is it when I'm storing a bunch of compressed video files? It's totally a waste at that point to even attempt to compress these files.
- treffer 4y agoThis is not how such storage systems work. If the whole storage system is for you and you only store compressed video files on it then perhaps. But we are talking about multi-tenant systems here. You will get better performance and/or better prices if the system finds compressible data for other customers. All the benefits hold, even for you with non compressible data, if there is an overall benefit. Hitting a bottleneck less often then this will reduce your tail latencies, too. It is just that _your files_ won't contribute to these improvements. But you get all the benefits as well.
- blibble 4y ago> They allow you to slice an arbitrary byte range out of an object which is technically harder to implement on a compressed file. this is pretty easy, you flush the compression buffer every megabyte or so and maintain an index maybe 50 lines of code
- jeffffff 4y agoSure, but now you've added an extra layer of indirection which can have a significant impact on performance
- klauspost 4y agoIt doesn't really have to impact performance. The index is generated easily as a side-effect of compression. And the index is only needed if you need to seek. I implemented this as part of the MinIO server. See "Seeking Compressed Files" here: https://blog.min.io/transparent-data-compression/ https://blog.min.io/transparent-data-compression/ We choose a compressor without literal compression for a faster baseline, but the concept remains the same.
- jeffffff 4y agoBut if you do need to seek, which is really common in data warehouse workloads for example, unless you keep the index in ram you have to do an extra IO on every seek to read the index
- blibble 4y agothere's always going to be some metadata for the file that needs to be looked up before you can start seeking (ACLs, sector/extent/cluster location, etc) the index goes in there, no extra seek needed
- jeffffff 4y agoyeah that isn't free either, it adds significant bloat to your metadata. with most enterprise customers encrypting and/or compressing data before putting it into s3, it doesn't seem like there would be much benefit. s3 really isn't the right layer to implement compression. filesystems aren't either. it's better to leave it up to the application.
- paulsutter 4y agoAmazon has millions of idle cpus available 24 hours a day (they can use all the idle time for all customer instances for whatever they want)
- eru 4y agoThat doesn't make it completely free. They still have opportunity costs.
- seabrookmx 4y agoIdle CPU's also use less power than ones running full tilt.
- Spooky23 4y agoI don’t work at AWS, but storage at scale is a funny beast, usually you’re constrained by IOPS, and if anything you have a surplus of CPU. If you can stuff more bits in an IO operation, you’re winning.
- natmaka 4y agoMoreover zstd quite unusual '--adapt' parameter enables it to "dynamically adapt compression level to perceived I/O conditions". Works for me (albeit the manpage states that "it can remain stuck at low speed when combined with multiple worker threads").
- deleted 4y ago[deleted]
- metadat 4y agoToo bad the flag doesn't come with detection for this environmental condition and then coordinate accordingly across processes.
- thecleaner 4y agoIs there a paper on how it "perceives" the I/O conditions?
- flaviut 4y agoI'd guess by using backpressure. Modify the compression level to try and keep the output buffer at 60% full.
- natmaka 4y agoIt reacts to the input buffer state (in bad I/O conditions it starves). On Linux PSI ( /proc/pressure/io ) probably provides more accurate information (the code already uses /proc/cpuinfo ). Detail: in the fileio.c module there are lines such as: if (oldIPos == inBuff.pos) inputBlocked++; /* input buffer is full and can't take any more : input speed is faster than consumption rate / if ( (inputBlocked > inputPresented / 8) / input is waiting often, because input buffers is full : compression or output too slow */ This impacts a 'speedChange' variable. Its potential values (an enum) are 'noChange', 'slower', and 'faster'. They are processed rather simply: if (speedChange == slower) { ((...)) compressionLevel ++; if (speedChange == faster) { ((...)) compressionLevel --;
- oogali 4y agoI think the point of the different storage tiers of AWS S3 is to get customers to classify their own data, then AWS can pick the right mix of hardware, software, and compute that satisfies AWS’s requirements for availability and COGS. If the difference between standard S3 and S3 Glacier was just slower disk, then rate limiting the customer would suffice. But if there’s a significant amount of compute thrown at data de-duplication, compression, and indexing, then it starts to clarify why there’s a pricing penalty for using Glacier with the same access patterns as one would use on standard storage.
- eru 4y ago> They charge by the GB and are not exactly super cheap so if the customer wants to store a big fat file of easily compressible zeros then whatever, they got their money. That's not a good argument: they could lower their costs with compression, still charge the same, and make more profit.
- naikrovek 4y agoI think an FPGA can probably compress quite a bit faster than a general purpose CPU, and an ASIC on a card which is network on one side and storage on the other could compress at line speed easily.