9 ms·
S3 Express Is All You Need
- BonoboIO 3y agoHas anyone here a usecase which would perform better with this new S3 Express Tier? And a second question, would it be worth the 8x times surcharge?
- parhamn 3y agoI think the key benefit brushed on by this article is the potential 10x improvement in access speeds (which has many applications, beyond reducing your s3 op charges). > S3 Express One Zone can improve data access speeds by 10x and reduce request costs by 50% compared to S3 Standard and scales to process millions of requests per minute.
- haimez 3y ago10x reduction in latency, higher storage costs with lower access costs (SSD instead of spinning disks). So high I/O, small files situations (with no need for cross AZ access) are where the benefits can be found.
- samtregar3 3y ago[dead]
- paulddraper 3y agoA cache with large blobs (images, etc)
- awoimbee 3y agoIf it's only a cache it should be on EBS, which is still way faster and 2x less expensive. I started a migration to s3 for such a project (container image caching) but then stopped when I realized what I was doing.
- paulddraper 3y ago1. You'd need an access/authentication layer on top of that. 2. Variable throughput may be a concern. 3. You may have availability concerns.
- rbranson 3y agoEBS attaches a single block storage volume to a single host[1]. S3 Express is a service-based object store. Apples and oranges. [1] Yes, I am aware of multi-attach but this introduces a scaling bottleneck and requires a fairly exotic setup.
- YetAnotherNick 3y agoYes, EBS is the gold standard but managing a EBS to scale up and down instantly, be available to multiple instances, lifecycle management, managing replica, switchover etc. are definitely not easy. And EBS are bad choice when throughput needed is very spiky.
- barsandtones 3y agoThis will work great with the s3 mount point that AWS recently released. This will outperform EFS if your application does not require full POSIX compatibility.
- maccard 3y agoI'm going to set up sccache [0] to use it tomorrow. We use MSVC, so EFS is off the cards. [0] https://github.com/mozilla/sccache/blob/main/docs/S3.md https://github.com/mozilla/sccache/blob/main/docs/S3.md
- tjoff 3y ago> However, the new storage class does open up an exciting new opportunity for all modern data infrastructure: the ability to tune an individual workload for low latency and higher cost or higher latency and lower cost with the exact same architecture and code. I get it, but at the same time that is also what you lost when you locked yourself in with a particular vendor.
- imheretolearn 3y ago> I get it, but at the same time that is also what you lost when you locked yourself in with a particular vendor. What are other viable practical alternative solution(s)?
- toomuchtodo 3y agoStorage adapter to talk S3 compatible to target, assuming you're not relying on vendor specific extensions or behavior (ie this). Off the top of my head, Backblaze B2, Cloudflare R2, etc are S3 compatible, and Minio locally. https://www.google.com/search?q=s3+compatible https://www.google.com/search?q=s3+compatible
- anamexis 3y agoThere are no vendor specific extensions or behavior here, are there? Isn't it just a different billing structure?
- jacobr1 3y agoNotifications, for event processing architectures aren't part of the API common to these systems
- williamdclt 3y agoI suppose “super low latency” is behaviour, in the sense that “a large enough quantitative difference is a qualitative difference”. If you rely on the perf and only S3 provides that, then you effectively are locked into S3 implementation
- Sirupsen 3y agoMost production storage systems/databases built on top of S3 spend a significant amount of effort building an SSD/memory caching tier to make them performant enough for production (e.g. on top of RocksDB). But it's not easy to keep it in sync with blob... Even with the cache, the cold query latency lower-bound to S3 is subject to ~50ms roundtrips [0]. To build a performant system, you have to tightly control roundtrips. S3 Express changes that equation dramatically, as S3 Express approaches HDD random read speeds (single-digit ms), so we can build production systems that don't need an SSD cache—just the zero-copy, deserialized in-memory cache. Many systems will probably continue to have an SSD cache (~100 us random reads), but now MVPs can be built without it, and cold query latency goes down dramatically. That's a big deal We're currently building a vector database on top of object storage, so this is extremely timely for us... I hope GCS ships this ASAP. [1] [0]: https://github.com/sirupsen/napkin-math https://github.com/sirupsen/napkin-math [1]: https://turbopuffer.com/ https://turbopuffer.com/
- jamesblonde 3y agoWe built HopsFS-S3 [0] for exactly this problem, and have running it as part of Hopsworks now for a number of years. It's a network-aware, write-through cache for S3 with a HDFS API. Metadata operations are performed on HopsFS, so you don't have the other problems list max listing operations return 1000 files/dirs. NVMe is what is changing the equation, not SSD. NVMe disks now have up to 8 GB/s, although the crap in the cloud providers barely goes to 2 GB/s - and only for expensive instances. So, instead of 40X better throughput than S3, we can get like 10X. Right now, these workloads are much better on-premises on the cheapest m.2 NVMe disks ($200 for 4TB with 4 GB/s read/write) backed by a S3 object store like Scality. [0] https://www.hopsworks.ai/post/faster-than-aws-s3 https://www.hopsworks.ai/post/faster-than-aws-s3
- dekhn 3y agothe numbers you're giving are throughput (byte/sec) not latency. The comment you reply to is talking mostly about latency - reporting that S3 object get latencies (time to open the object and return its head) in the single-digits ms, where S3 was 50ms before. BTW EBS can do 4GB/sec per volume. But you will pay for it.
- throwitaway222 3y agoI don't understand why EFS never gets major shout outs - it's way better than S3: systems can mount it as a drive, shared across systems, already has had super low latency... Not sure what s3 express is really useful for if EFS already exists.
- candiddevmike 3y agoEFS is really expensive and has terrible latency with small files in my experience
- brazzledazzle 3y agoYeah the main reason is that it's incredibly expensive. You can improve performance by allocating ahead of time but NFS has never been at its best when working with a bunch of tiny files.
- richieartoul 3y agoDo you have any more details you can share about the performance of EFS? I've never met anyone who has actually used it in anger.
- gchamonlive 3y agoThroughput scales with the amount of data in it, it is in the docs. So depending on the application, even if latency is better, the speeds are atrocious at lower volumes of persisted data.
- saddlerustle 3y agoThat’s not true anymore with EFS Elastic Throughput
- a2tech 3y agoYes, I built a moderately large system on it that used lots of small shared files. The performance was fairly terrible. There's weird little niggles with it--we had random slowdowns, throughput issues, and things just didn't work quite right. It was an ok solution for what we were doing, but several times I came really close to just dumping it and standing up an NFS server using EBS volumes. I also used it a couple of times to store webroots and that was a complete disaster with systems that had lots of small files (Drupal I'm looking at you).
- emgeee 3y agosome additional context here is that warpstream is building a Kakfa compatible streaming system that uses s3 as the object store. This allows them to leverage cheap zone transfer costs for redundancy + automatic storage tiering to cut down on the costs of running and maintaining these systems. This has previously come at the cost of latency due to s3's read/write speeds but with S3 this makes them more competitive with Confluent Kafka's managed offerings for these latency sensitive applications. IMO warpstream is a really cool product and this new S3 offering makes them even better
- refset 3y agoI am eager to hear how it will affect their latency numbers: > Engineering is about trade-offs, and we’ve made a significant one with WarpStream: latency. The current implementation has a P99 of ~400ms for Produce requests because we never acknowledge data until it has been durably persisted in S3 and committed to our cloud control plane. In addition, our current P99 latency of data end-to-end from producer-to-consumer is around 1s via https://www.warpstream.com/blog/kafka-is-dead-long-live-kafka https://www.warpstream.com/blog/kafka-is-dead-long-live-kafk...
- fswd 3y agoI solved this problem locally. When uploading a file to the server before going to S3 it is cached in redis. Whenever the codebase needs to use the file, it checks redis, and if it is not there it fetches it and caches it again.
- jamiesonbecker 3y agoExactly. Write-through cache is exactly how Userify[0] used to work for self-hosted versions. (when it was Python, we used Redis to keep state synced across multiple processes, but now that it's a Go app, we do all the caching and state management in memory using Ristretto[1]) However, we now install by default to local disk filesystem, since it's much faster to just do a periodic S3 hot sync, like with restic or aws-cli, than to treat S3 as the primary backing store, or just version the EBS or instance volume. The other reason you might want to use S3 as a primary is if you use a lot of disk, but our files are compressed and extremely small, even for a large installation with tens of thousands of users and instances. 0. https://userify.com https://userify.com (ssh key management + sudo for teams) 1. https://github.com/dgraph-io/ristretto https://github.com/dgraph-io/ristretto
- avinassh 3y agoWhat were the reasons to move from Redis to Ristretto? Both seem to be very different, since Redis is distributed where as Ristretto is local to the process.
- jamiesonbecker 3y agoIn our case, Python (because of the GIL) required us to have a single python process per core in order to take advantage of multiple cores, and so we needed Redis to maintain a unified memory state across all the cores, but Go can automatically span across multiple cores. We also saw about a 10x speedup by moving all caching into the server process, and since it was all in the same process, we no longer had to compress and encrypt data before sending to Redis. We still checkpoint the moving server state, encrypted and compressed, to disk every sixty seconds, just like Redis would do with BGSAVE, so we can start back up within a few seconds (actually faster than the old Redis after a restart.)
- osti 3y agoIf I'm not wrong, this is the low latency S3 that is written in Rust. Finally launched after years in the making.
- FridgeSeal 3y agoDo you have any sources for that? Very interested to know more about this.
- osti 3y agoUnfortunately I don't, this is already internal information that I don't know if I should say here. I never worked on S3 and I no longer work at AWS so someone from within would have to weigh in.
- paulddraper 3y agoSurely being written in a non-Rust language is not responsible for an extra 40ms of latency, right? Or is rust really that magic?
- osti 3y agoOf course not, it's designed differently from the original S3. AWS came out with this to compete with Azure premium blob storage, which has very good first byte latency, and Azure had it 4 years ago.. https://azure.microsoft.com/en-us/blog/premium-block-blob-storage-a-new-level-of-performance/ https://azure.microsoft.com/en-us/blog/premium-block-blob-st...
- estebarb 3y agoShardStore? (More info: https://www.thestack.technology/aws-shardstore-s3/ https://www.thestack.technology/aws-shardstore-s3/ ) it seems that it was deployed years ago.
- osti 3y agoNah, that looks very different, one of the stated goals of S3 Express is to minimize latency, which is the only thing about the Rust S3 that I remember.
- francoismassot 3y agoWe tested S3 Express for our search engine quickwit [0] a couple of weeks ago. While this was really satisfying on the performance side, we were a bit disappointed by the price, and I mostly agree with the article on this matter. I can see some very specific use cases where the pricing should be OK but currently, I would say most of our users will just stay on the classic S3 and add some local SSD caching if they have a lot of requests. [0] https://github.com/quickwit-oss/quickwit/ https://github.com/quickwit-oss/quickwit/
- kernelsanderz 3y agoI'd be fascinated if you could share your insights from using this. Where does the pricing fall down? And is the latency/throughput a big improvement for this use case? (ie. externalizing a search index).
- fulmicoton 3y agoI ran the benchmark at Quickwit. I confirm it works as intended. I was extremely excited about this feature, primarily interested in the decreased GET request cost, and secondly the lower latency. Unfortunately the price model puts it in a place where it is the right technology only for some very rare places. In a nutshell the key thing you need to know is: - The storage is 6.4x expensive than classic S3. - The GET requests are 2x cheaper (with additional cost for large requests). - Your data is replicated within a single region. - latency is single digit ms. From a pure cost wise point of view, the realm where it makes sense to use it is there, but small, and often competes more with EBS than it competes with S3.
- kernelsanderz 3y agothank so much for sharing. Amazing product BTW!
- Hixon10 3y agoDid you have access to their private preview version, or it is some another S3 Express, not S3 Express One Zone?
- mgaunard 3y agoMany S3 implementations appear to simply be transparent downloads to disk rather than a true "use the network as a disk".
- promocha 3y ago> “Of course the AWS S3 Express storage costs are still 8x higher than S3 standard, but that’s a non issue for any modern data storage system. Data can be trivially landed into low latency S3 Express buckets, and then compacted out to S3 Standard buckets asynchronously. Most modern data systems already have a form of compaction anyways, so this “storage tiering” is effectively free.” This is key insight. The data storage cost essentially becomes negligible and latency goes down by a magnitude by making S3 Express as a buffer storage then moving data to standard S3. I see a future where most data-intensive apps would use S3 as main storage layer.
- tomjakubowski 3y agosounds a bit like CPU caches and main memory
- haimez 3y agoOr like SSD’s vs spinning disks…
- otabdeveloper4 3y agoDid you conveniently ignore egress costs?
- kristianp 3y agoI saw "X is all you Need" with the "Attention is all you need" paper [1], which launched the Transformer upon the world. Is it the first instance of that phrase? [1] https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762
- collinc777 3y agoWill this improve running sqlite on s3?
- api 3y ago“All You Need Considered Harmful” - most cliche title?