20 ms·
Amazon S3 Adds Put-If-Match (Compare-and-Swap)
- throwaway314155 2y ago[flagged]
- ramon156 2y agoHonestly if it was fast and uninvasive, I wouldn't mind it at all
- earth2mars 2y agoWhat is stopping you not doing it now? I know Q is not good (hallucinates, slow, requires sign in) But it's wise to explain what your gripe is about than saying which you can always do.
- throwaway314155 2y agoMy gripe was with the Explainer modal that covers the entire article upon visiting the site.
- koolba 2y agoThis combined with the read-after-write consistency guarantee is a perfect building block (pun intended) for incremental append only storage atop an object store. It solves the biggest problem with coordinating multiple writers to a WAL.
- IgorPartola 2y agoRename for objects and “directories” also. Atomic.
- ncruces 2y agoBoth this and read-after-write consistency is single object. So coordinating writes to multiple objects still requires… creativity.
- sillysaurusx 2y agoFinally. GCP has had this for a long time. Years ago I was surprised S3 didn’t.
- ncruces 2y agoGCS is just missing x-amz-copy-source-range in my book. Can we have this Google? … Please?
- mannyv 2y agoGCP still doesn't have triggers out of beta last time i checked (which was a while ago).
- seansmccullough 2y agoAzure Storage has also had this for years - https://learn.microsoft.com/en-us/rest/api/storageservices/specifying-conditional-headers-for-blob-service-operations https://learn.microsoft.com/en-us/rest/api/storageservices/s...
- 1a527dd5 2y agoBe still my beating heart. I have lived to see this day. Genuinely, we've wanted this for ages and we got half way there with strong consistency.
- ncruces 2y agoMight finally be possible to do this on S3: https://pkg.go.dev/github.com/ncruces/go-gcp/gmutex https://pkg.go.dev/github.com/ncruces/go-gcp/gmutex
- phrotoma 2y agoHuh. Does this mean that the AWS terraform provider could implement state locking without the need for a DDB table the way the GCP provider does?
- rekwah 2y agoI started looking into this but DeleteObject doesn't support these conditional headers on general purpose buckets; only directory buckets (Express Zone One).
- paulddraper 2y agoSo....given CAP, which one did they give up
- johnrob 2y agoI’d wager that the algorithm is slightly eager to throw a consistency error if it’s unable to verify across partitions. Since the caller is naturally ready for this error, it’s likely not a problem. So in short it’s the P :)
- tonymet 2y agogood example of how a simple feature on the surface (a header comparison) requires tremendous complexity and capacity on the backend.
- offmycloud 2y agoIf the default ETag algorithm for non-encrypted, non-multipart uploads in AWS is a plain MD5 hash, is this subject to failure for object data with MD5 collisions? I'm thinking of a situation in which an application assumes that different (possibly adversarial) user-provided data will always generate a different ETag.
- revnode 2y agoMD5 hash collisions are unlikely to happen at random. The defect was that you can make it happen purposefully, making it useless for security.
- aphantastic 2y agoSure, but theoretically you could have a system where a distributed log of user generated content is built via this CAS//MD5 primitive. A malicious actor could craft the data such that entries are dropped.
- revnode 2y agoMy understanding of the feature, and correct me if I'm wrong, is that you are not granted write access based on a hash. You already have write access. You can use the hash to avoid overwriting someone else's data that was appended to the file in between you checking the file and writing to it. If you already have write access, the hash is irrelevant. As a bad actor, you can corrupt the data without it. MD5 should not be used for anything security related. Granting write access based on an MD5 hash would be a huge no-no.
- aphantastic 2y agoRight, the issue comes when a trusted writer is logging data that is sourced from an untrusted party. Imagine a transaction log being a blob per-customer with many lines corresponding to price, sku, etc, that additionally have some “memo” field provided by the customer. A trusted distributed worker process is responsible for taking incoming requests by the user, pulling their blob down, appending the line based on the request, and CAS’ing it back in (retrying on failure). With enough effort, a particularly devious user could issue many requests with ‘memo’s engineered to not alter the MD5 of their log. This would cause some lines to be lost. An audit of their account transaction log would be unable to accurately reflect the requests they made to the service, and the failure would be invisible. This is obviously a bit contrived – I’ll be the first to admit. But if the incentives were to exist for this to be worth someone’s time for some system, I think it would be likely to see it come up eventually.
- gravitronic 2y agoFirst thing I thought when I saw the headline was "oh! I should tell Sirupsen"
- JoshTriplett 2y agoIt's also possible to enforce the use of conditional writes: https://aws.amazon.com/about-aws/whats-new/2024/11/amazon-s3-enforcement-conditional-write-operations-general-purpose-buckets/ https://aws.amazon.com/about-aws/whats-new/2024/11/amazon-s3... My biggest wishlist item for S3 is the ability to enforce that an object is named with a name that matches its hash. (With a modern hash considered secure, not MD5 or SHA1, though it isn't supported for those either.) That would make it much easier to build content-addressible storage.
- cmeacham98 2y agoIs there any reason you can't enforce that restriction on your side? Or are you saying you want S3 to automatically set the name for you based on the hash?
- JoshTriplett 2y ago> Is there any reason you can't enforce that restriction on your side? I'd like to set IAM permissions for a role, so that that role can add objects to the content-addressible store, but only if their name matches the hash of their content. > Or are you saying you want S3 to automatically set the name for you based on the hash? I'm happy to name the files myself, if I can get S3 to enforce that. But sure, if it were easier, I'd be thrilled to have S3 name the files by hash, and/or support retrieving files by hash.
- mdavidn 2y agoI think you can presign PutObject calls that validate a particular SHA-256 checksum. An API endpoint, e.g. in a Lambda, can effectively enforce this rule. It unfortunately won’t work on multipart uploads except on individual parts.
- UltraSane 2y agoThe hash of multipart uploads is simply the hash of all the part hashes. I've been able to replicate it.
- Sirupsen 2y agoTo avoid any dependencies other than object storage, we've been making use of this in our database (turbopuffer.com) for consensus and concurrency control since day one. Been waiting for this since the day we launched on Google Cloud Storage ~1 year ago. Our bet that S3 would get it in a reasonable time-frame worked out! https://turbopuffer.com/blog/turbopuffer https://turbopuffer.com/blog/turbopuffer
- amazingamazing 2y agoInteresting that what’s basically an ad is the top comment - it’s not like this is open source or anything - can’t even use it immediately (you have to apply for access). Totally proprietary. At least elasticsearch is APGL, saying nothing of open search which also supports use of S3
- viraptor 2y agoSomeone made an informed technical bet that worked out. Sounds like HN material to me. (Also, is it really a useful ad if you can't easily use the product?)
- amazingamazing 2y agoWorked out how? There’s no implementation. It’s just conjecture.
- CubsFan1060 2y agoI feel dumb for asking this, but can someone explain why this is such a big deal? I’m not quite sure I am grokking it yet.
- deleted 2y ago[deleted]
- Sirupsen 2y agoThe short of it is that building a database on top of object storage has generally required a complicated, distributed system for consensus/metadata. CAS makes it possible to build these big data systems without any other dependencies. This is a win for simplicity and reliability.
- CubsFan1060 2y agoThanks! Do they mention when the comparison is done? Is it before, after, or during an upload? (For instance, if I have a 4tb file in a multi part upload, would I only know it would fail as soon as the whole file is uploaded?)
- poincaredisk 2y agoI imagine, for it to make sense, that the comparison is done at the last possible moment, before atomically swapping the file contents.
- Nevermark 2y agoI can imagine it might be useful to make this a choice for databases with high frequency small swaps and occasional large ones. 1) default, load-compare-&-swap for small fast load/swaps. 2) optional, compare-load-&-swap to allow a large load to pass its compare, and cut in front of all the fast small swap that would otherwise create an un-hittable moving target during its long loads for its own compare. 3) If the load itself was stable relative to the compare, then it could be pre-loaded and swapped into a holding location, followed by as many fast compare-&-swaps as needed to get it into the right location.
- rrr_oh_man 2y agoCould anybody explain for the uninitiated?
- msoad 2y agoIt ensures that when you try to upload (or “put”) a new version of a file, the operation only succeeds if the file on the server still has the exact version (ETag) you specify. If someone else has updated the file in the meantime, your upload is blocked to prevent overwriting their changes. This is especially useful in scenarios where multiple users or processes are working on the same data, as it helps maintain consistency and avoids accidental overwrites. This is using the same mechanism as HTTP's `If-None-Match` header so it's easier to implement/learn
- rrr_oh_man 2y agoThank you! That was extremely helpful (and written in a way that is easy to understand)!
- wanderingmind 2y agoDoes this mean, in theory we will be able to manage multiple concurrent writes/updates to s3 without having to use new solutions like Regatta[1] that was recently launched? https://news.ycombinator.com/item?id=42174204 https://news.ycombinator.com/item?id=42174204
- huntaub 2y agoHere's how I would think about this. Regatta isn't the best way to add synchronization primitives to S3, if you're already using the S3 API and able to change your code. Regatta is most useful when you need a local disk, or a higher performance version of S3. In this case, the addition of these new primitives actually just makes Regatta work better for our customers -- because we get to achieve even stronger consistency.
- dvektor 2y ago[rejected] error: failed to push some refs to remote repository Finally we can have this with s3 :)
- mdaniel 2y agoRelevant: https://github.com/awslabs/git-remote-s3#readme https://github.com/awslabs/git-remote-s3#readme https://news.ycombinator.com/item?id=41887004 https://news.ycombinator.com/item?id=41887004
- vlovich123 2y agoI implemented that extension in R2 at launch IIRC. Thanks for catching up & helping move distributed storage applications a meaningful step forward. Intended sincerely. I'm sure adding this was non-trivial for a complex legacy codebase like that.
- ipython 2y agoI can't wait to see what abomination Cory Quinn can come up with now given this new primitive! (see previous work abusing Route53 as a database: https://www.lastweekinaws.com/blog/route-53-amazons-premier-database/ https://www.lastweekinaws.com/blog/route-53-amazons-premier-...)
- stevefan1999 2y agoSo...are we closer to getting to use S3 as a...you guessed it...a database? With CAS, we are probably able to get a basic level of atomicity, and S3 itself is pretty durable, now we have to deal with consistency and isolation...although S3 branded itself as "eventually consistent"...
- mr_toad 2y agoPeople who want all those features use something like Delta Lake on top of object storage.
- User23 2y agoThere was a great deal of interest in gossip protocols, eventual consistency, and such at Amazon in the mid oughts. So much so that they hired a certain Cornell professor along with the better part of his grad students to build out those technologies.
- gynther 2y agoS3 is strongly consistent since 4 years ago. https://aws.amazon.com/blogs/aws/amazon-s3-update-strong-read-after-write-consistency/ https://aws.amazon.com/blogs/aws/amazon-s3-update-strong-rea...
- amazingamazing 2y agoIronically with this and lambda you could make a serverless sqlite by mapping pages to objects, using http range reads to read the db and lambda to translate queries to the writes in the appropriate pages via cas. Prior to this it would require a server to handle concurrent writers, making the whole thing a nonstarter for “serverless”. Too bad performance would be terrible without a caching layer (ebs).
- captn3m0 2y agoFor read heavy workloads, you could cache the results at cloudfront. Maybe we will someday see Wordpress-on-Lambda-to-Sqlite-over-S3.
- m_d_ 2y agos3fs's https://github.com/fsspec/s3fs/pull/917 https://github.com/fsspec/s3fs/pull/917 was in response to the IfNoneMatch feature from the summer. How would people imagine this new feature being surfaced in a filesystem abstraction?
- grahamj 2y agobender_neat.gif
- maglite77 2y agoNoting that Azure Blob storage supports e-tag / optimistic controls as well (via If-Match conditions)[1], how does this differ? Or is it the same feature? [1]: https://learn.microsoft.com/en-us/azure/storage/blobs/concurrency-manage https://learn.microsoft.com/en-us/azure/storage/blobs/concur...
- simonw 2y agoIt's the same feature. Google Cloud Storage has it too: https://cloud.google.com/storage/docs/request-preconditions#precondition_criteria https://cloud.google.com/storage/docs/request-preconditions#...
- paulsutter 2y agoWhat’s amazing is that it took them so long to add these functions
- thayne 2y agoNow if only you had more control over the ETag, so you could use a sha256 of the total file (even for multi-part uploads), or a version counter, or a global counter from an external system, or a logical hash of the content as opposed to a hash of the bytes.
- serbrech 2y agoWhy is standard etag support making the frontpage?
- vytautask 2y agoAn open-source implementation of Amazon S3 - MinIO has had it for almost two years (relevant post: https://blog.min.io/leading-the-way-minios-conditional-write-feature-for-modern-data-workloads/ https://blog.min.io/leading-the-way-minios-conditional-write...). Strangely, Amazon is catching up just now.
- topspin 2y agoThat's not "strange" to me. Object storage has been a long time coming, and it's still being figured out: the entirely typical process of discovering useful and feasible primitives that expand applicability to more sophisticated problems. This is obviously going occur first in smaller and/or younger, more agile implementations, whereas AWS has the problem of implementing this at pretty much the largest conceivable scale with zero risk. The lag is, therefore, entirely unsurprising.
- aseipp 2y agoIt's not surprising at all. The scale of AWS, in particular S3, is nearly unfathomable, and the kind of solutions they need for "simple" things are totally different at that size. S3 was doing 1.1million requests a second back in 2013.[1] I wouldn't be surprised if they saw over 100mil/req/sec globally by now. That's 100 million requests a second that need strong read-your-write consistency and atomicity at global scale. The number of pieces they had to move into place for this to happen is probably quite the engineering tale. [1] https://aws.amazon.com/blogs/aws/amazon-s3-two-trillion-objects-11-million-requests-second https://aws.amazon.com/blogs/aws/amazon-s3-two-trillion-obje...
- lttlrck 2y agoIsn't this compare-and-set rather than compare-and-swap?
- torginus 2y agoAh so its not only me that uses AWS primitives for hackily implementing all sorts of synchronization primitives. My other favorite pattern is implementing a pool of workers by quering ec2 instances with a certain tag in a stopped state and starting them. Starting the instance can succeed only once - that means I managed to snatch the machine. If it fails, I try again, grabbing another one. This is one of those things that I never advertised out of professional shame, but it works, its bulletproof and dead simple and does not require additional infra to work.
- _zoltan_ 2y agothis actually sounds interesting. do you precreate the workers beforehand and then just keep them in a stopped state?
- torginus 2y agoyeah. one of the goals was startup time, so It made sense to precreate them. In practice we never ran out of free machines (and if we did, I have a cdk script to make more), and inifnite scaling is a pain in the butt anyways due to having to manage subnets etc. Cost-wise we're only paying for the EBS volumes for the stopped instances which are like 4GB each, so they cost practically nothing, we spend less than a dollar per month for the whole bunch.
- zild3d 2y agoWarm pools are a supported feature in AWS on auto scaling groups. Works as you're describing (have a pool of instances in stopped state ready to use, only pay for EBS volume if relevant) https://aws.amazon.com/blogs/compute/scaling-your-applications-faster-with-ec2-auto-scaling-warm-pools/ https://aws.amazon.com/blogs/compute/scaling-your-applicatio...
- zerd 2y agoIf you want even fast startups restarting stopped instances is apparently faster https://depot.dev/blog/faster-ec2-boot-time https://depot.dev/blog/faster-ec2-boot-time
- anonymousDan 2y agoWould be interesting to understand how they've implemented it and they whether there is any perf impact on other API calls.
- londons_explore 2y agoSo we can now implement S3-as-RAM for a worldwide million-core linux VM?
- juggli 2y agofinally
- spprashant 2y agoI had no idea people rely on S3 beyond dumb storage. It almost feels like people are trying to build out a distributed OLAP database in the reverse direction.
- amne 2y ago1. SELECT ... INTO OUTFILE S3 2. glue jobs to partition by some columns reporting uses 3. query with athena 4. ??? 5. profit (celebrate reduced cost) This thing costs couple $ a month for ~500gb of data. Snowflake wanted crazy amounts of money for the same thing.