7 ms·
A race condition in Aurora RDS
- deleted 11mo ago[deleted]
- redwood 11mo agoA good reminder of how people developing a mental model of adding read replicas as a way to scale is a slippery slope. At the end of the day you're scaling only one specific part of your system with certain consistency dynamics that are difficult to reason about
- terminalshort 11mo agoWorks fine for workloads like: 1. I need to grab some rows from a table 2. Eventual consistency is good enough And that's a lot of workloads.
- candiddevmike 11mo agoAs a user, I've come to realize the situations where I think eventual consistency (or delayed processing) are good enough aren't the same as the folks developing most products. Nothing annoys me more than stuff not showing up immediately or having to manually refresh.
- darth_avocado 11mo agoSometimes users want everything to show up immediately, but not pay extra for the feature. Everything real time is expensive. Eventual consistency is a good thing for most systems.
- terminalshort 11mo agoFor a workload where you need true read after write you can just send those reads to the writer. But even if you don't there are plenty of workarounds here. You can send a success response to the user when the transaction commits to the writer and update the UI on response. The only case where this will fail is if the user manually reloads the page within the replication lag window and the request goes to the reader. This should be exceedingly rare in a single region cluster, and maybe a little less rare in a multi-region set up, but still pretty rare. I almost never see > 1s replication lag between regions in my Aurora clusters. There are certainly DB workloads where this will not be true, but if you are in a high replication lag cluster, you just don't want to use that for this type of UI dependency in the first place.
- nilamo 11mo agoI think the key here is just proper notifications. Yes it's eventually consistent, but having a "processing" or "update in progress" is a huge improvement over showing a user old data.
- redwood 11mo agoThe future you or future team member may struggle to reason about that in the future
- morshu9001 11mo agoThat's readonly. RW workloads usually don't tolerate eventual consistency on the thing they're writing.
- terminalshort 11mo agoYeah, if you have a mix of reads and writes in a workflow, you gotta hit the writer node. But a lot of times an endpoint is only reading data from a particular DB.
- nijave 11mo agoYou can hit the same problems horizontally scaling compute. One instance reads from the DB, a request hits a different instance which updates the DB. The original instance writes to the DB and overwrites the changes or makes decisions based on stale data. More broadly a distributed system problem
- gtowey 11mo agoThis article seems to indicate that manually triggered failovers will always fail if your application tries to maintain its normal write traffic during that process. Not that I'm discounting the author's experience, but something doesn't quite add up: - How is it possible that other users of Aurora aren't experiencing this issue basically all the time? How could AWS not know it exists? - If they know, how is this not an urgent P0 issue for AWS? This seems like the most basic of basic usability features is 100% broken. - Is there something more nuanced to the failure case here such as does this depend on transactions in-progress? I can see how maybe the failover is waiting for in-flight transactions to close and then maybe hits a timeout where it proceeds with the other part of the failover by accident. That could explain why it doesn't seem like the issue is more widespread.
- maherbeg 11mo agoYeah I agree, this seems like a pretty critical feature to the Aurora product itself. We saw a similar behavior recently where we had a connection pooler in between which indicates something wrong with how they propagate DNS changes during the failover. wtf aws
- CaptainKanuk 11mo agoWhenever we have to do any type of AWS Aurora or RDS cluster modification in prod we always have the entire emergency response crew standing by right outside the door. Their docs are not good and things frequently don't behave how you expect them to.
- ekropotin 11mo agoOh, well, it’s always DNS!
- dboreham 11mo agoAlthough the article has an SEO-optimized vibe, I think it's reasonable to take it as true until refuted. My rule of thumb is that any rarely executed, very tricky operation (e.g. database writer fail over) is likely to not work because there are too many variables in play and way too few opportunities to find and fix bugs. So the overall story sounds very plausible to me. It has a feel of: it doesn't work under continuous heavy write load, in combination with some set of hardware performance parameters that plays badly with some arbitrary time out. Note that the system didn't actually fail. It just didn't process the fail over operation. It reverted to the original configuration and afaics preserved data.
- jansommer 11mo agoPeople who have experience with Aurora and RDS Postgres: What's your experience in terms of performance? If you dont need multi A-Z and quick failover, can you achieve better performance with RDS and e.g. gp3 64.000 iops and 3125 throughput (assuming everything else can deliver that and cpu/mem isn't the bottleneck)? Aurora seems to be especially slow for inserts and also quite expensive compared to what I get with RDS when I estimate things in the calculator. And what's the story on read performance for Aurora vs RDS? There's an abundance of benchmarks showing Aurora is better in terms of performance but they leave out so much about their RDS config that I'm having a hard time believing them.
- shawabawa3 11mo ago> 3125 throughput Max throughput on gp3 was recently increased to 2GB/s, is there some way I don't know about of getting 3.125?
- jansommer 11mo agoThis is super confusing. Check out the RDS Postgres calculator with gp3: > General Purpose SSD (gp3) - Throughput > gp3 supports a max of 4000 MiBps per volume But the docs say 2000. Then there's IOPS... The calculator allows up to 64.000 but on [0], if you expand "Higher performance and throughout" it says > Customers looking for higher performance can scale up to 80,000 IOPS and 2,000 MiBps for an additional fee. [0] https://aws.amazon.com/ebs/general-purpose/ https://aws.amazon.com/ebs/general-purpose/
- nijave 11mo agoRDS PG stripes multiple gp3 volumes so that's why RDS throughput is higher than gp3 I think 80k IOPs on gp3 is a newer release so presumably AWS hasn't updated RDS from the old max of 64k. iirc it took a while before gp3 and io2 were even available for RDS after they were released as EBS options Edit: Presumably it takes some time to do testing/optimizations to make sure their RDS config can achieve the same performance as EBS. Sometimes there are limitations with instance generations/types that also impact whether you can hit maximum advertised throughput
- grhmc 11mo agoYikes! This is exactly the kind of invariant I'd expect Aurora to maintain on my behalf. It is why I pay them so much...
- dangoodmanUT 11mo agoIt did, the storage layer did not allow for concurrent writes.
- bob1029 11mo ago> Aurora's architecture differs from traditional PostgreSQL in a crucial way: it separates compute from storage. I find this approach very compelling. MSSQL has a similar thing with their hyperscale offering. It's probably the only service in Azure that I would actually use.
- robinduckett 11mo agoGlad to know I’m not crazy.
- theanomaly 11mo agoAWS Support initially pushed back and suggested it's because of high replication lag but they were looking at metrics that were more than 24 hours old. What kind of failure did you encounter? I really want to understand what edge case we triggered in their failover process - especially since we could not reproduce it in other regions.
- robben1234 11mo agoMy cluster recently started to failover every few days whenever it experiences the load to trigger scale up from 1-2 to 20+ acu. And then I also encountered errors just like op in my app layer about trying to execute a write query via read-only transaction. The workaround so far is to invalidate connection on error. When app reconnects the cluster write endpoint correctly leads to current primary.
- d1egoaz 11mo ago> AWS has indicated a fix is on their roadmap, but as of now, the recommended mitigation aligns with our solution: use Aurora’s Failover feature on an as-needed basis and ensure that no writes are executed against the DB during the failover. Is there a case number where we can reach out to AWS regarding this recommendation?
- paranoidrobot 11mo agoYeah. I'd like this too. We use Aurora MySQL but I would like to be able to point to that and ask if it applies to us.
- time0ut 11mo agoWow. This is alarming. We have done a similar operation routinely on databases under pretty write intensive workloads (like 10s of thousands of inserts per second). It is so routine we have automation to adjust to planned changes in volume and do so a dozen times a month or so. It has been very robust for us. Our apps are designed for it and use AWS’s JDBC wrapper. Just one more thing to worry about I guess…
- dangoodmanUT 11mo agoNot really: Their storage layer worked perfectly and prevented the ACID violations.
- almosthere 11mo agoprobably should have added postgres to end of title
- evanelias 11mo agoAbsolutely this. The differences between Aurora Postgres and Aurora MySQL are quite significant. A failover bug affecting one doesn't imply the same bug exists in the other. A lot of people seem to have the misconception that "Aurora" is its own unique database system, with different front-ends "pretending" to be Postgres or MySQL, but that isn't the case at all.
- ldkge 11mo agoAm I the only one who misread that as “AI race condition”?
- dangoodmanUT 11mo agoThis confirms a lot of what their engineers preach: The lego brick model. They made the storage layer in total isolation, and they made sure that it guaranteed correctness for exclusive writer access. When the upstream service failed to also make its own guarantees, the data layer was still protected. Good job AWS engineering!
- halifaxbeard 11mo agoI think OP is wrong in their hypothesis based on the logs they share and the root cause AWS support provided them. I think the promotion fails to happen and then an external watchdog notices that it didn’t, and kills everything ASAP as it’s a cluster state mismatch. The message about the storage subsystem going away is after the other Postgres process was kill -9’d.
- YouAreWRONGtoo 11mo ago[dead]
- halfmatthalfcat 11mo agoCC pm. MgtzkskskzjauHjhffd
- shayonj 11mo agoSadly, its not the first time I have noticed unexpected and odd behaviors from Aurora PostgreSQL offering. I noticed another interesting (and still unconfirmed) bug with Aurora PostgreSQL around their Zero Downtime Patching. During an Aurora minor version upgrade, Aurora preserves sessions across the engine restart, but it appears to also preserve stale per-session execution state (including the internal statement timer). After ZDP, I’ve seen very simple queries (e.g. a single-row lookup via Rails/ActiveRecord) fail with `PG::QueryCanceled: ERROR: canceling statement due to statement timeout` in far less than the configured statement_timeout (GUC), and only in the brief window right after ZDP completes. My working theory is that when the client reconnects (e.g. via PG::Connection#reset), Aurora routes the new TCP connection back to a preserved session whose “statement start time” wasn’t properly reset, so the new query inherits an old timer and gets canceled almost immediately even though it’s not long-running at all.