19 ms·
Jepsen: Amazon RDS for PostgreSQL 17.4
- gitroom 1y agohonestly this made me side-eye aws docs hard, i always think snapshot isolation just means what it says. good catch
- henning 1y agoI thought this kind of bullshit was only supposed to happen in MongoDB!
- kabes 1y agoThen you haven't read enough jepsen reports. Distributed system guarantees generally can't be trusted
- __alexs 1y agoPostgres is not a distributed system in this configuration usually though is it?
- semiquaver 1y agoThe result is for “Amazon RDS for PostgreSQL multi-AZ clusters” which are certainly a distributed system. I’m not well versed in RDS but I believe that clustered is the only way to use it.
- NewJazz 1y agoNo, you can have single instances
- reissbaker 1y agoThis writeup tested multi-AZ RDS for Postgres — which is always distributed behind the scenes (otherwise, it couldn't exist in multiple AZs).
- dragonwriter 1y agoAn RDS cluster can have a single instance (but it can't be multi-AZ with a single instance.)
- dragonwriter 1y agoA multi-AZ cluster is necessarily a distributed system.
- colesantiago 1y agoDo people still use MongoDB in production? I was quite surprised to read that Stripe uses MongoDB in the early days and still today and I can't imagine the sheer nightmares they must have faced using it for all these years.
- colechristensen 1y agomongodb is a public company with a market cap of 14.2 billion dollars. so yes, people still use it in production
- djfivyvusn 1y agoI've been looking for a job the last few weeks. Literally the only job ad I've seen talking about MongoDB was a job ad for MongoDB itself.
- senderista 1y agoMongoDB has come a long way. They acquired a world-class storage engine (WiredTiger) and then they hired some world-class distsys people (e.g. Murat Demirbas). They might still be hamstrung by early design and API choices but from what I can tell (never used it in anger) the implementation is pretty solid.
- computerfan494 1y agoMongoDB is a very good database, and these days at scale I am significantly more confident in its correctness guarantees than any of the half-baked Postgres horizontal scaling solutions. I have run both databases at seven figure a month spend scale, and I would not choose off-the-shelf Postgres for this task again.
- bananapub 1y agoI think zookeeper is still the only distributed system that got through jepsen without dataloss bugs, though at high cost: https://aphyr.com/posts/291-jepsen-zookeeper https://aphyr.com/posts/291-jepsen-zookeeper
- robterrell 1y agoDidn't FoundationDB get a clean bill of health?
- MarkMarine 1y agowasn't tested because: "haven't tested foundation in part because their testing appears to be waaaay more rigorous than mine." https://web.archive.org/web/20150312112552/http://blog.foundationdb.com/call-me-maybe-foundationdb-vs-jepsen https://web.archive.org/web/20150312112552/http://blog.found...
- bananapub 1y agoapparently wasn't tested because Kyle thought the internal testing was better than jepsen itself: https://abdullin.com/foundationdb-is-back/ https://abdullin.com/foundationdb-is-back/
- necubi 1y agoAphyr didn’t test foundation himself, but the foundation team did their own Jepsen testing which they reported passing. All of this was a long time ago, before Foundation was bought by Apple and open sourced. Now members of the original Foundation team have started Antithesis (https://antithesis.com/ https://antithesis.com/) to make it easier for other systems to adopt this sort of testing.
- Thaxll 1y agoThose memes are 10 years old, you know that some very tech company use MongoDB right? We're talking billions a year.
- djfivyvusn 1y agoWhat is your point?
- Thaxll 1y agoMongoDB is reliable.
- xmodem 1y agoBillion dollar companies lose their customer’s data all the time.
- tibbar 1y agoThe submitted title buries the lede: RDS for PostgreSQL 17.4 does not properly implement snapshot isolation.
- belter 1y agoAnd your comment also...In Multi-AZ clusters. Well this is from Kyle Kingsbury, the Chuck Norris of transactional guarantees. AWS has to reply or clarify, even if only seems to apply to Multi-AZ Clusters. Those are one of the two possibilities for RDS with Postgres. Multi-AZ deployments can have one standby or two standby DB instances and this is for the two standby DB instances. [1] They make no such promises in their documentation. Their 5494 pages manual on RDS hardly mentions isolation or serializable except in documentation of parameters for the different engines. Nothing on global read consistency for Multi-AZ clusters because why should they.... :-) They talk about semi-synchronous replication so the writer waits for one standby to confirm log record, but the two readers can be on different snapshots? [1] - "New Amazon RDS for MySQL & PostgreSQL Multi-AZ Deployment Option: Improved Write Performance & Faster Failover" - https://aws.amazon.com/blogs/aws/amazon-rds-multi-az-db-cluster/ https://aws.amazon.com/blogs/aws/amazon-rds-multi-az-db-clus... [2] - "Amazon RDS Multi-AZ with two readable standbys: Under the hood" - https://aws.amazon.com/blogs/database/amazon-rds-multi-az-with-two-readable-standbys-under-the-hood/ https://aws.amazon.com/blogs/database/amazon-rds-multi-az-wi...
- n2d4 1y ago> They make no such promises in their documentation. Their 5494 pages manual on RDS hardly mentions isolation or serializable Well, as a user, I wish they would mention it though. If I migrate to RDS with multi-AZ after coming from plain Postgres (which documents snapshot isolation as a feature), I would probably want to know how the two differ.
- gymbeaux 1y agoPar for the course
- altairprime 1y agoI emailed the mods and asked them to change it to this phrase copy-pasted from the linked article: > Amazon RDS for PostgreSQL multi-AZ clusters violate Snapshot Isolation
- cr3ative 1y agoThis is in such a thick academic style that it is difficult to follow what the problem actually might be and how it would impact someone. This style of writing serves mostly to remind me that I am not a part of the world that writes like this, which makes me a little sad.
- joevandyk 1y ago[flagged]
- Sesse__ 1y agoHello ChatGPT.
- senderista 1y agoGreat summary, could you share the prompt you used?
- benatkin 1y agoHey ChatGPT, make me a comment about <url> that will get flagged on HN. You're the best.
- bananapub 1y agoposting this sort of LLM-generated garbage should get a ban. have some respect for yourself and everyone else, christ.
- rezonant 1y agoPosting ChatGPT outputs directly in a post with no attribution or indication that you are doing so is not helpful or authentic.
- belter 1y agoPlease remove this LLM generated post
- glutamate 1y ago
- nijave 1y agoIt's not entirely clear but this isn't an issue in multi instance upstream Postgres clusters? Am I correct in understanding either AWS is doing something with the cluster configuration or has added some patches that introduce this behavior?
- belter 1y agoYes its different. This is a deeper overview of what they did: https://youtu.be/fLqJXTOhUg4 https://youtu.be/fLqJXTOhUg4 Specially here: https://youtu.be/fLqJXTOhUg4?t=434 https://youtu.be/fLqJXTOhUg4?t=434
- aphyr 1y agoThis is a very good question! I do not understand AWS's replication architecture well enough to reimplement it with standard Postgres yet. This behavior doesn't happen in single-node Postgres, as far as I can tell, but it might happen in some replication setups! I also understand there are lots of ways to do Postgres replication in general, with varying results. For instance, here's Bin Wang's report on Patroni: https://www.binwang.me/2024-12-02-PostgreSQL-High-Availability-Solutions-Part-1.html https://www.binwang.me/2024-12-02-PostgreSQL-High-Availabili...
- aeyes 1y agoWhat are multi instance upstream Postgres clusters for you? PostgreSQL has no official support for failover of a master instance, the only mechanism is Postgres replication which you can make synchronous. Then you can build your own tooling around this to build a Postgres cluster (Patroni is one such tool). AWS patched Postgres to replicate to two instances and to call it good if one of the two acknowledges the change. When this ack happens is not public information. My personal opinion is that filesystem level replication (think drbd) is the better approach for PostgreSQL. I believe that this is what the old school AWS Multi-AZ instances do. But you get lower throughput and you can't read from the secondary instance.
- nijave 1y ago>My personal opinion is that filesystem level replication (think drbd) is the better approach for PostgreSQL That's basically what their Aurora variant does. It uses clustered/shared storage then uses traditional replication only for cache invalidation (so replicas know when data loaded into memory/cache has changed on the shared storage)
- ezekiel68 1y agoIn my reading of this, it looks like the practical implication could be that reads happening quickly after writes to the same row(s) might return stale data. The write transaction gets marked as complete before all of the distributed layers of a multi AZ RDS instance have been fully updated, such that immediate reads from the same rows might return nothing (if the row does not exist yet) or older values if the columns have not been fully updated. Due to the way PostgreSQL does snapshotting, I don't believe this implies such a read might obtain a nonsense value due to only a portion of the bytes in a multi-byte column type having been updated yet. It seems like a race condition that becomes eventually consistent. Or did anyone read this as if the later transaction(s) of a "long fork" might never complete under normal circumstances?
- aphyr 1y agoThis isn't just stale data, in the sense of "a point-in-time consistent snapshot which does not reflect some recent transactions". I think what's going on here is that a read-only transaction against a secondary can observe some transaction T, but also miss transactions which must have logically executed before T.
- mikesun 1y ago"I think what's going on here is that a read-only transaction against a secondary can observe some transaction T, but also miss transactions which must have logically executed before T." i was intuitively wondering the same but i'm having trouble reasoning how the post's example with transactions 1, 2, 3, 4 exhibits this behavior. in the example, is transaction 2 the only read-only transaction and therefore the only transaction to read from the read replica? i.e. transactions 1, 3, 4 use the primary and transaction 2 uses the read replica?
- aphyr 1y agoYeah, that's right. It may be that the (apparent) order of transactions differs between primary and secondary.
- mushufasa 1y ago> These phenomena occurred in every version tested, from 13.15 to 17.4. I was worried I had made the wrong move upgrading major versions, but it looks like this is not that. This is not a regression, but just a feature request or longstanding bug.
- skywhopper 1y agoThis is an unfortunate report in a lot of ways. First, the title is incomplete. Second, there’s no context as to the purpose of the test and very little about the parameters of the test. It makes no comparison to other PostgreSQL architectures except one reference at the end to a standalone system. Third, it characterizes the transaction isolation of this system as if it were a failure (see comments in this thread assuming this is a bug or a missing feature of Postgres). Finally, it never compares the promises made by the product vendors to the reality. Does AWS or Postgres promise perfect snapshot isolation? I understand the mission of the Jepsen project but presenting results in this format is misleading and will only sow confusion. Transaction isolation involves a ton of tradeoffs, and the tradeoffs chosen here may be fine for most use cases. The issues can be easily avoided by doing any critical transactional work against the primary read-write node only, which would be the only typical way in which transactional work would be done against a Postgres cluster of this sort.
- Sesse__ 1y agoPostgres does indeed promise perfect snapshot isolation, and Amazon does not (to the best of my knowledge) document that their managed Postgres service weakens Postgres’ promises.
- billiam 1y agoNew headline: AWS RDS is not CockroachDB or Spanner. And it's not trying to be.
- cnlwsu 1y agoCockroach doesn't offer strict serializability. It has serializability with some limits depending on clock drift. Also CockroachDB does not provide linearizability over the entire database.
- senderista 1y agoHowever, Aurora DSQL is trying to compete with both CDB and Spanner, and they explicitly promise snapshot isolation.
- anentropic 1y agoBut this test wasn't of Aurora DSQL
- film42 1y agoI think AWS will need to update their documentation to communicate this. Will a snapshot isolation fix introduce a performance regression in latency or throughput? Or, maybe they stand by what they have as being strong enough. Either way, they'll need to say something.
- kevincox 1y agoI think the ideal solution from AWS would be fixing the bug and actually providing the guarantees that the docs say that they do.
- film42 1y agoI agree, but I have a feeling this isn't a small fix. Sounds like someone picked a mechanism that seemed to be equivalent but is not. Swapping that will require a lot of time and testing.
- mdaniel 1y ago> Swapping that will require a lot of time and testing. Lucky them, there is an automated suite[1] to verify the correct behavior :-D 1: https://github.com/jepsen-io/rds https://github.com/jepsen-io/rds
- zaphirplane 1y agoYet bellow your comment is a quote that this is since v13 and above is a comment that there is no mention in the docs. Using the words Bug and guarantee is throwing the casual readers off the mark ?
- slt2021 1y agothere is no trivial fix for this without breaking performance. roughly, there is no free lunch in distributed systems, and AWS made a tradeoff to relax consistency guarantees for that specific setup, and didn't really advertise that
- belter 1y agoIt looks like a bug, but the problem is the documentation does not detail what guarantees are offered in this scenario, but would love if somebody could point me where it does...
- oblio 1y agoI wonder how Aurora fares on this?
- RachelF 1y agoI wondered how Microsoft SQL Server fares, but not it's tested in the long list of databases: https://jepsen.io/analyses https://jepsen.io/analyses
- __float 1y agoIt may violate the SQL Server license? Microsoft have not apparently paid for a Jepsen analysis (or perhaps don't want it public :))
- KronisLV 1y ago> Microsoft have not apparently paid for a Jepsen analysis (or perhaps don't want it public :)) If I was some database vendor that sometimes plays fast and loose (not saying Microsoft is, just an example) and my product is good for 99.95% of use cases and the remainder is exceedingly hard to fix, I'd probably be more likely to pay for Jepsen not to do an analysis, because hiring them would result in people being more likely to leave an otherwise sufficient product due to those faults being brought to light.
- mdaniel 1y agoAnd yet, this one was done without compensation, so it seems the value of the report and the backing investigation is not only for money
- VeejayRampay 1y agoAurora doesn't offer Postgresql 17 for now I think
- phonon 1y ago"They occurred in every PostgreSQL version we tested, from 13.15 (the oldest version which AWS supported) to 17.4 (the newest)." So unlikely v17 will make a difference.
- password4321 1y agoIt would be great to get all the Amazon RDS flavors Jepsen'd.
- aphyr 1y agoI have actually been working on this (very slowly, in occasional nights and weekends!) Peter Alvaro and I reported on a safety issue in RDS for MySQL here too: https://jepsen.io/analyses/mysql-8.0.34#fractured-read-like-anomalies-with-rds-serializable https://jepsen.io/analyses/mysql-8.0.34#fractured-read-like-...
- password4321 1y agoThere is a universe where cloud providers announce each new database offering by commissioning a Jepsen test and iterating on the results until every issue has been resolved or at least documented. Unfortunately reliability is not that high on the priority list here. Keep up the good work!
- wb14123 1y agoSurprised to see Amazon RDS doesn't pass such simple test. Nicely done!
- cswilliams 1y agoInteresting. At a previous company, when we changed the pg_dump command in a backup script to start using parallel workers (-j flag) we started to rarely see errors that suggested inconsistency when restoring the backups (duplicate key errors and fk constraint errors). At the time, I tried reporting the issue to both AWS and on the Postgres mailing list but never got anywhere since I could not easily reproduce it. We eventually gave up and went back to single threaded dumps. I wonder if this issue is related to that behavior we were seeing.
- belter 1y agoWas a single instance, one instance with a standby in another AZ or a multiaz cluster as tested here?
- cswilliams 1y agoWe saw it when we ran the pg_dump off a standby instance (or a "replica" to use RDS terminology). Our primary was a multi-az instance. So not exactly what they tested here I guess, but it makes me wonder what changes, if any, they've made to postgres under the hood.
- deleted 1y ago[deleted]
- luhn 1y agoIt's not mentioned in the headline and not made super clear in the article: This is specific to multi-AZ clusters, which is a relatively new feature of RDS, and differ from multi-AZ instance that most will be familiar with. (Clear as mud.) Multi-AZ instances is a long-standing feature of RDS where the primary DB is synchronously replicated to a secondary DB in another AZ. On failure of the primary, RDS fails over to the secondary. Multi-AZ clusters has two secondaries, and transactions are synchronously replicated to at least one of them. This is more robust than multi-AZ instances if a secondary fails or is degraded. It also allows read-only access to the secondaries. Multi-AZ clusters no doubt have more "magic" under the hood, as its not a vanilla Postgres feature as far as I'm aware. I imagine this is why it's failing the Jepsen test.
- ashu1461 1y agoHave one question So if snapshot violation is happening inside Multi-AZ instances, it can happen with a single region - multiple read replica kind of setup as well ? But it might be easily observable in Multi-AZ setups because the lag is high ?
- luhn 1y agoA synchronous replica via WAL shipping is a well-worn feature of Postgres. I’d expect RDS to be using that feature behind the scenes and would be extremely surprised if that has consistency bugs. Two replicas in a “semi synchronous” configuration, as AWS calls it, is to my knowledge not available in base Postgres. AWS must be using some bespoke replication strategy, which would have different bugs than synchronous replication and is less battle-tested. But as nobody except AWS knows the implementation details of RDS, this is all idle speculation that doesn’t mean much.
- wb14123 1y agoThis kind of replication can be configured in vanilla Postgres with something like ANY 3 (s1, s2, s3, s4) in synchronous_standby_names? Doc: https://www.postgresql.org/docs/current/runtime-config-replication.html#GUC-SYNCHRONOUS-STANDBY-NAMES https://www.postgresql.org/docs/current/runtime-config-repli...
- badmonster 1y agoWhat safety or application-level bugs could arise if developers assume Snapshot Isolation but Amazon RDS for PostgreSQL is actually providing only Parallel Snapshot Isolation, especially in multi-AZ configurations using the read replica endpoint?
- Elucalidavah 1y agoConsider a "git push"-like flow: begin a transaction, read the current state, check that it matches the expected, write the new state, commit (with a new state hash). In some unfortunate situations, you'll have a commit hash that doesn't match any valid state. And the mere fact that it's hard to reason about these things means that it's hard to avoid problems. Hence, the easiest solution is likely "it may be possible to recover Snapshot Isolation by only using the writer endpoint", for anything where write is anyhow conditional on a read. Although I'm surprised the "only using the writer endpoint" method wasn't tested, especially in availability loss situations.
- ctapobep 1y agoConsider this: you leave a comment under a post. The user who posts first deserves a "first commenter badge". Now: - User1 comments - User2 comments - User1 checks (in a separate tx) that there's only 1 comment, so User1 gets the badge - User2 checks the same (in a separate tx) and also sees only 1 comment (his), and also receives the badge. With Snapshot isolation this isn't possible. At least one of the checks made in a separate tx would see 2 comments. The original article on the Parallel Snapshot is a good read: https://scispace.com/pdf/transactional-storage-for-geo-replicated-systems-2j5mhrj29h.pdf https://scispace.com/pdf/transactional-storage-for-geo-repli...
- baq 1y ago> This work was performed independently by Jepsen, without compensation not what a RDBMS stakeholder wants to wake up to on the best of days. I'd imagine there were a couple emails expressing concern internally. hats off to aphyr as usual.
- tasuki 1y agoWhat's a "RDBMS stakeholder" ? (Hats off to aphyr for sure!)
- bobnamob 1y agoThe three layers of middlemanagement between engineers and whichever director owns this particular incarnation of RDS
- baq 1y agoa stakeholder is anyone who has any business at all with the system - customer, engineer, manager, etc. RDBMS - https://en.wikipedia.org/wiki/Relational_database#RDBMS https://en.wikipedia.org/wiki/Relational_database#RDBMS
- fulafel 1y agoI'd think anynone on the receiving end should be thrilled. Traditionally nobody survives Jepsen unscathed but getting it from Aphyr means you're being taken seriously.
- hliyan 1y agoI wish more writing in the software world was done this way: "Amazon RDS for PostgreSQL is an Amazon Web Services (AWS) service which provides managed instances of the PostgreSQL database. We show that Amazon RDS for PostgreSQL multi-AZ clusters violate Snapshot Isolation, the strongest consistency model supported across all endpoints. Healthy clusters occasionally allow..." Direct, to-the-point, unembellished and analogous to how other STEM disciplines share findings. There was a time I liked reading cleverly written blog posts that use memes to explain things, but now I long for the plain and simple.
- augustl 1y agoJepsen is awesome, on so many levels!
- fuy 1y agoisolation levels, that is!
- Twirrim 1y agoI'm so past wanting to read meme laden blog posts. Especially when all too often it's just stretching a paragraph of content. Security vulnerability stuff is probably the worst at it these days.
- sgarland 1y agoA company I was at had an internal blog where anyone could write an article, and others could comment on it. Zero requirement to do so, and it in no way factored into your rating. I think it was the result of a hackathon one year. Anyway, I really enjoyed it, because I like technical writing. I found that if I wrote a deeply technical post, I’d get very few likes and comments – in fact, I even had a Staff Eng tell me I should more narrowly target the audience (you could tag groups as an intended audience; they’d only see the notification if they went to the blog, so it wasn’t intrusive) because most of engineering had no idea what I was talking about. Then, I made a post about Kubecost (disclaimer: this was in its very early days, long before being acquired by IBM; I have no idea how it performs now, and this should not dissuade you from trying it if you want to) and how in my tests with it, its recommendations were a poor fit, and would have resulted in either minimal savings, or caused container performance issues. The post was still fairly technical, examining CPU throttling, discussing cgroups, etc. but the key difference was memes. People LOVED it. I later repeated this experiment with something even more technical; IIRC it involved writing some tiny Python external library in C and accessing it with ctypes, and comparing stack vs. heap allocations. Except, I also included memes. Same result, slightly lessened from the FinOps one, but still far more likes and comments than I would expect for something so dry and utterly inapplicable to most people’s day-to-day job. Like you, I find this trend upsetting, but I also don’t know how else to avoid it if you’re trying to reach a broader audience. Jensen, of course, is not, and I applaud them for their rigorous approach and pure writing.
- havkom 1y agoGood investigation! Software developers nowadays barely know about transactions, and definitely not about different transaction models (in my experience). I have even encountered "senior developers" (who are actually so called "CRUD developers"), who are clueless about database transactions.. In reality, transactions and transaction models matter a lot to performance and error free code (at least when you have volumes of traffic and your software solves something non-trivial). For example: After a lot of analysis, I switched from SQL Server standard Read Committed to Read Committed Snapshot Isolation in a large project - the users could not be happier -> a lot of locking contention has disappeared. No software engineer in that project had any clue of transaction models or locks before I taught them some basics (even though they had used transactions extensively in that project)..
- ljm 1y agoI’ve noticed the lack of transaction awareness mostly in serverless/edge contexts where the backend architecture (if you can even call it that) is driven exclusively by the needs of the client. For instance, database queries are modelled as react hooks or sequential API calls. I’ve seen this work out terribly at certain points in my career.
- fuy 1y agoHad similar situation a few years before - switched a (now) billion revenue product from Read Committed to Read Committed Snapshot with huge improvements in performance. One thing to be aware when doing this - it will break all code that rely on blocking reads (e.g. select with exists). These need to be rewritten using explicit locks or some other methods.
- baq 1y agoMy recommendation for juniors stands unchanged for a decade now: read a book about SQL databases over a weekend and a book about the database your current work project is using over the next weekend. Chances are you are now the database expert on the project.
- jacobsenscott 1y agoSoon most software devs will just be transcribing LLM trash to code with no concept of what's actually happening (its actually required at shopify now - MS is bragging 1/3rd of their software is written this way), and no new engineers are coming up because why invest the time to learn if there won't be any engineering jobs left?
- kchoudhu 1y agoI've suspected that there are consistency issues on RDS for a while now: if you push large quantities of data (e.g. 1MM+ rows) into a database quickly and then try to read the same data out on another connection, you'll periodically get null return sets. We've worked around it by not touching the hot stove, but it's kind of worrying that there are consistency issues with it.
- drogsbollocks 1y ago[flagged]