10 ms·
Upgrading Uber's MySQL Fleet
- sinaptia_dev 2y ago[flagged]
- vivzkestrel 2y agoAnyone has any ideas why Uber doesn't use PostgreSQL?
- hu3 2y agohttps://eng.uber.com/postgres-to-mysql-migration/ https://eng.uber.com/postgres-to-mysql-migration/
- cyberax 2y agoThey switched from PG to MySQL because they need to update highly concurrent tables, and PG created tons of bloat as a result. MySQL uses locking instead of optimistic concurrency.
- drastic_fred 2y agoWow! Less than 200qps per node, quite a redundancy. They also mentioned its read heavy. With aws dynamo db it would be at least 10 times cheaper for their given workload. Although, i do not know how much data they store and serve or how they are additionally (analytics?, etc) using this fleet for.
- edf13 2y ago3 million queries/second across 16k nodes seems pretty heavy on redundancy?
- withinboredom 2y agoThat's 200 qps per node, assuming perfect load balancing.
- sgarland 2y agoI was going to say, that's absolutely nothing. They state 2.1K clusters and 16K nodes; if you divide those, assuming even distribution, you get 7.6 instances/cluster. Round down because they probably rounded up for the article, so 1 primary and 6 replicas per cluster. That's still only ~1400 QPS / cluster, which isn't much at all. I'd be interested to hear if my assumptions were wrong, or if their schema and/or queries make this more intense than it seems.
- pgwhalen 2y ago> assuming even distribution I don't work for Uber, but this is almost certainly the assumption that is wrong. I doubt there is just a single workload duplicated 2.1K times. Additionally, different regions likely have different load.
- 620gelato 2y ago2100 clusters, 16k nodes, and data is replicated across every node "within a cluster" with nodes placed in different data centers/regions. That doesn't sound unreasonable, on average. But I suspect the distribution is likely pretty uneven.
- tiffanyh 2y agoWhy upgrade to v8.0 (old LTS) and not v8.4 (current LTS)? Especially given that end-of-support is only 18-months from now (April 2026) … when end-of-support of v5.7 is what drive them to upgrade in the first place. https://en.m.wikipedia.org/wiki/MySQL https://en.m.wikipedia.org/wiki/MySQL
- gostsamo 2y ago> Several compelling factors drove our decision to transition from MySQL v5.7 to v8.0: Edit: for the downvoters, the parent comment was initially a question.
- johannes1234321 2y agoSince direct upgrade to 8.4 isn't supported. They got to go to 8.0 first. Also: 8.0 is old and most issues have been found. 8.4 probably has more unknowns.
- tallanvor 2y agoAccording to the article they started the project in 2023. Given that 8.4 was released in April 2024, that wasn't even an option when they started.
- hu3 2y agoThe upgrade initiative started somewhere in 2023 according to the article. MySQL 8.4 was released in April 30, 2024. Their criteria for a "battle tested" MySQL version is probably much more rigorous than the average CRUD shop.
- paulryanrogers 2y agoConsidering several versions of 8.0 had a crashing bug if you renamed a table, waiting is probably the right choice.
- blindriver 2y agoYou’re not renaming tables when you’re at scale.
- whalesalad 2y agoSo satisfying to do a huge upgrade like this and then see the actual proof in the pudding with all the reduced latencies and query times.
- hu3 2y agoYeah some numbers caught my attention like ~94% reduction in overall database lock time. And to think they never have to worry about VACUUM. Ahh the peace.
- anonzzzies 2y agoYeah, until vacuum is gone, i'm not touching postgres. So many bad experiences with our use cases over the decades. I guess most people don't have our uses, but i'm thinking Uber does.
- RedShift1 2y agoMaybe just vacuum much more aggressively? Also there have been a lot of changes to the vacuuming and auto vacuuming process these last few years, you can pretty much forget about it.
- anonzzzies 2y agoNot in our experience; for our cases it is still a resource hog. We discussed it even less than a year ago with core devs and with a large postgres consultancy place; they said postgres doesn't fit our use case which was already our conclusion, no matter how much we want it to be. Mysql is smooth as butter. I have nothing to win from picking mysql just that it works; I rather use postgres as features / not oracle but... Edit; also, as can be seen here in responses, and elsewhere on the web when discussing this, the fans say it's no problem, but many less religious users feel it's a massive design flaw (perfectly logical at the time, not so logical now) that sometimes will stop users from using it, which is a shame
- 2y ago
- candiddevmike 2y agoDoes Uber still use Docstore? I'd imagine having built an effectively custom DB on top of MySQL made this upgrade somewhat inconsequential for most apps.
- dravikontor1212 2y ago[dead]
- geitir 2y agoYes
- gregoriol 2y agoWait until they find out they have to upgrade to 8.4 now
- gregoriol 2y agoAnd also all the passwords away from mysql_native_password
- sgarland 2y agoThey've got until 9.0 for that, it just gives deprecation warnings in 8.4.
- evanelias 2y agoMore specifically, mysql_native_password is disabled by default in 8.4, but can be re-enabled if needed: https://www.skeema.io/blog/2024/05/14/mysql84-surprises/#authentication-changes https://www.skeema.io/blog/2024/05/14/mysql84-surprises/#aut...
- gregoriol 2y agoDeprecation warnings are in 8.0. It's disabled in 8.4. If you are up-to-date with all your libraries it all should go well, but if some project is stuck on some old code, mostly old mysql libraries, one might get surprises when doing the switch away.
- johannes1234321 2y agoWhich one should do anyways. mysql_native_password is considered broken for a out ten years. (Broken for people who can access the hashed form of the password on the server)
- rafram 2y agoDid they have ChatGPT (re)write this? The writing style is very easy to identify, and it’s grating.
- deleted 2y ago[deleted]
- OsrsNeedsf2P 2y ago> The writing style is very easy to identify, Really? At n=1 the rate seems to be 0
- indulona 2y ago[flagged]
- kavunr 2y agoI had a similar reaction when reading https://engineeringblog.yelp.com/2024/10/migrating-from-postgres-to-mysql.html https://engineeringblog.yelp.com/2024/10/migrating-from-post...
- BarryMilo 2y agoI get the urge to standardize infrastructure but... wow. After reading the whole thing, I get why they thought it was worth doing, but you'd think they'd just hire a Postgres guy or two. Especially when they looked at the missing features, this feels like a downgrade...
- kccqzy 2y agoIn big companies they never just "hire a guy or two"; the bus factor would be atrocious in this case, not to mention vacations and such. The minimum unit of hiring is one team, about five people five or take. So the difference they are looking at is either hiring a team of Postgres people or forcing everyone on one existing team to learn Postgres deeply. From this perspective, standardizing on infrastructure makes more sense now.
- indulona 2y agoeveryone is praising postgres and shitting on mysql, until performance matters. then, when all these big tech companies turn to mysql, the postgres fanboys cry in unison.
- lenerdenator 2y agoThere'd be less of that if there weren't a feeling that MySQL is the high-grade free hit that Oracle uses to get you hooked on their low-grade crap.
- m4r1k 2y agoUber's collaboration with Percona is pretty neat. The fact that they've scaled their operations without relying on Oracle's support is a testament to the expertise and vision of their SRE and SWE teams. Respect!
- tiffanyh 2y agoAren't they using Persona in lieu of Oracle. So it's kind of the same difference, no?
- paulryanrogers 2y agoWord on the street is Oracle contracts are expensive and hard to cancel, like a deal with the devil. Not sure if their MySQL support is any different than Oracle DB itself
- paradite 2y agoI can tell from a mile away that this is written by ChatGPT / Claude, at least partially. "This distinction played a crucial role in our upgrade planning and execution strategy." "Navigating Challenges in the MySQL Upgrade Journey" "Finally, minimizing manual intervention during the upgrade process was crucial."
- mannyv 2y agoOnce ChatGPT puts in "we did the needful" we're all doomed.
- greenchair 2y agoDear sir, we are having a P1 incident, Prashant please revert.
- brunocvcunha 2y agoI can tell just by the frequency of the word “delve”
- notinmykernel 2y agoAgree. Repetition (e.g., crucial) in ChatGPT is an issue.
- deleted 2y ago[deleted]
- traceroute66 2y ago> I can tell from a mile away that this is written by ChatGPT / Claude, at least partially. Whilst it may smell of ChatGPT/Claude, I think the answer is actually simpler. Look at the authors of the blog, search LinkedIn. They are all based in India, mostly Bangalore. It is therefore more likely to be Indian English. To be absolutely clear, for absolute avoidance of doubt: This is NOT intended a racist comment. Indians clearly speak English fluently. But the style and flow of English is different. Just like it is for US English, Australian English or any other English. I am not remotely saying one English is better than another ! If, like me, you have spent many hours on the phone to Bangalore call-centres, you will recognise many of the stylistic patterns present in the blog text.
- John23832 2y agoAnyone else get a "Not Acceptable" response?
- internetter 2y agoI did but it worked on a private tab
- menaerus 2y agoLately there's been a shitload of sponsored $$$ and anti-MySQL articles so it's kinda entertaining that their authors are being slapped in their face by Uber, completely unintended.
- greenie_beans 2y agoyall should prioritize your focus so you can do better at vetting drivers who don't almost kill me
- lenerdenator 2y agoTheir focus is prioritized according to what returns maximum value to their shareholders.
- greenie_beans 2y agobeep boop i'm a capitalist robot pretty sure safe travels is critical to maximum value to their shareholders (aka stfu or tell me how this blog post has anything to do with maximize shareholder value https://www.uber.com/en-JO/blog/upgrading-ubers-mysql-fleet/#h-motivation-for-the-upgrade https://www.uber.com/en-JO/blog/upgrading-ubers-mysql-fleet/... ... shareholder value is a dumb ass thing to prioritize over human life)
- lenerdenator 2y ago> pretty sure safe travels is critical to maximum value to their shareholders Well, it is... sort of. Obviously you can't have Uber be a guaranteed way to be robbed by a highwayman, but when you've cleared out most taxis in a given city, you can start to dictate the terms by which customers accept your service. And if that means including language in your ToS that shove your customers into a binding arbitration agreement [0] that effectively shield you from the risks of hiring incompetent or malicious drivers, well... that's what that means. [0]https://www.npr.org/2024/10/02/nx-s1-5136615/uber-car-crash-lawsuit-uber-eats-arbitration-terms https://www.npr.org/2024/10/02/nx-s1-5136615/uber-car-crash-...
- greenie_beans 2y agoyou can rationalize your way around the issue all day but it still don't make it right.
- 2y ago
- jauntywundrkind 2y agoHaving spent a couple months doing a corporate mandated password rotation on our services - a number of which weren't really designed for password rotation - happy to see the dual password thing mentioned. Being able to load in a new password while the current one is active is where it's at! Trying to coordinate a big bang where everyone flips over at the same time is misery, and I spent a bunch of time updating services to not have to do that! Great enhancement. I wonder what other datastores have dual (or more) password capabilities?
- johannes1234321 2y agoI can't answer with an overview on who got such a feature, but "every" system got a different way of doing that: rotating usernames as well. Create a new user with new password. This isn't 100% equal as ownership (thus permissions with DEFINER) in stored procedures etc. needs some thought, but bad access using outdated username is simpler to trace (as username can be logged etc. contrary to passwords; while MySQL allows for tracing using performance_schema logging incl. user defined connection attributes which may ease finding the "bad" application)
- donatj 2y agoInterestingly we just went through basically the same upgrade just a couple days ago for similar reasons. We run Amazon Aurora MySQL and Amazon is finally forcing us to upgrade to 8.0. We ended up spinning up a secondary fleet and bin log replicating from our 5.7 master to the to-be 8.0 master until everything made the switch over. I was frankly surprised it worked, but it did. It went really smoothly.
- takeda 2y agoAFAIK the 8.0 release is one where Oracle breaks compatibility. So anyone considering MariaDB needs to switch before going to 8.0, otherwise switching will be much more painful.
- paulryanrogers 2y agoIIRC it wasn't a big break from 5.7 unless you used some odd grouping or relied on its old+lossy defaults. Though the single threaded performance fell off a cliff, tanking my CI performance. Maria DB evolved very differently. I'm not sure how they stack up performance wise.
- donatj 2y agoYou can re-enable the old group by behavior, which we did.
- sandGorgon 2y agoso how does an architecture like "2100 clusters" work. so the write apis will go to a database that contains their data ? how is this done - like a user would have history, payments, etc. are all of them colocated in one cluster ? (which means the sharding is based on userid) ? is there then a database router service that routes the db query to the correct database ?
- ericbarrett 2y agoA query for a given item goes to a router*, as you said, that directs it to a given shard which holds the data. I don't know Uber's schema, but usually the data is "denormalized" and you are not doing too many JOINs etc. Probably a caching layer in front as well. If you think this sounds more like a job for a K/V store than a relational database, well, you'd be right; this is why e.g. Facebook moved to MyRocks. But MySQL/InnoDB does a decent job and gives you features like write guarantees, transactions, and solid replication, with low write latency and no RAFT or similar nondeterministic/geographically limited protocols. * You can also structure your data so that the shard is encoded in the lookup key so the "routing" is handled locally. Depends on your setup
- bob1029 2y agoI imagine it works just like any multi-tenant SaaS product wherein you have a database per customer (region/city) with a unified web portal. The primary difference being that this is B2C and the ratio of customers per database is much greater than 1.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- denysonique 2y agoWhy didn't they move to MariaDB instead? A faster than MySQL 8 drop-in replacement.
- evanelias 2y agoWhile it is indeed often faster, it isn't drop-in. MySQL and MariaDB have diverged over the years, and each has some interesting features that the other lacks. I wrote a summary of the DDL / table design differences between MySQL and MariaDB, and that topic alone is fairly long: https://www.skeema.io/blog/2023/05/10/mysql-vs-mariadb-schema/ https://www.skeema.io/blog/2023/05/10/mysql-vs-mariadb-schem... Another area with major differences is replication, especially when moving beyond basic async topologies.
- aorth 2y agoWow, I hadn't realized that MySQL and MariaDB diverged so much! In the last year I've started seeing some prominent applications like Apache Superset and Apache AirFlow claiming they don't support—or even test on—MariaDB at all.
- dweekly 2y agoAm I the only one who saw "delve" at the top of the article and immediately thought "ah, an AI generated piece"? Well, that and the over-structured components of the analysis with nearly uniform word count per point and high-complexity but low signal-to-noise vocabulary using phraseology not common to the domain being discussed. (The article doesn't scan as written by an SRE/DBA.)
- devbas 2y agoThe introduction seems to have AI sprinkled all over it: ..we embarked on a significant journey, ..in this monumental upgrade.
- deleted 2y ago[deleted]
- xyst 2y agoI wonder if an upgrade like this would be less painful if the db layer was containerized? The migration process they described would be less painful with k8s. Especially with 2100+ nodes/VMs
- BowBun 2y agoA pipe dream. Having recently interacted with a modern k8s operator for Postgres, it lacked support for many features that had been around for a long time. I'd be surprised if MySQL's operators are that much better. Also consider the data layer, which is going to need to be solved regardless. Of course at Uber's scale they could write their own, I guess. At that point, if you're reaching in and scripting your pods to do what you want, you lose a lot of the benefits of convention and reusability that k8s promotes.
- jcgl 2y ago> it lacked support for many features that had been around for a long time Care to elaborate at all? Were they more like missing edge cases or absent core functionality? Not to imply that missing edge cases aren’t important when it comes to DB ops.
- __turbobrew__ 2y agoI can tell you that k8s starts to have issues once you get over 10k nodes in a single cluster. There has been some work in 1.31 to improve scalability but I would say past 5k nodes things no longer “just work”: https://kubernetes.io/blog/2024/08/15/consistent-read-from-cache-beta/ https://kubernetes.io/blog/2024/08/15/consistent-read-from-c... The current bottleneck appears to be etcd, boltdb is just a crappy data store. I would really like to try replacing boltdb with something like sqlite or rocksdb as the data persistence layer in etcd but that is non-trivial. You also start seeing issues where certain k8s operators do not scale either, for example cilium cannot scale past 5k nodes currently. There are fundamental design issues where the cilium daemonset memory usage scales with the number of pods/endpoints in the cluster. In large clusters the cilium daemonset can be using multiple gigabytes of ram on every node in your cluster. https://docs.cilium.io/en/stable/operations/performance/scalability/report/ https://docs.cilium.io/en/stable/operations/performance/scal... Anyways, the TL;DR is that at this scale (16k nodes) it is hard to run k8s.
- remon 2y agoImpressive numbers at a glance but that boils down to ~140qps which is between one and two orders of magnitude below what you'd expect a normal MySQL node typically would serve. Obviously average execution time is mostly a function of the complexity of the query but based on Uber's business I can't really see what sort of non-normative queries they'd run at volume (e.g. for their customer facing apps). Uber's infra runs on Amazon AWS afaik and even taking some level of volume discount into account they're burning many millions of USD on some combination of overcapacity or suboptimal querying/caching strategies.
- aseipp 2y agoDividing the fleet QPS by the number of nodes is completely meaningless because it assumes that queries are distributed evenly across every part of the system and that every part of the system is uniform (e.g. it is unclear what the read/write patterns are, proportion of these nodes are read replicas or hot standbys, if their sizing and configuration are the same). That isn't realistic at all. I would guess it is extremely likely that hot subsets of these clusters, depending on the use case, see anywhere from 1 to 4 orders of magnitude higher QPS than your guess, probably on a near constant basis. Don't get me wrong, a lot of people have talked about Uber doing overengineering in weird ways, maybe they're even completely right. But being like "Well, obviously x/y = z, and z is rather small, therefore it's not impressive, isn't this obvious?" is the computer programming equivalent of the "econ 101 student says supply and demand explain everything" phenomenon. It's not an accurate characterization of the system at all and falls prey to the very thing you're alluding to ("this is obvious.")
- 0cf8612b2e1e 2y agoSimple enough just to think about localities and time of day. New York during Tuesday rush hour could be more load than all of North Dakota sees in a month. Even busy cities probably drop down to nothing on a weekday at 3am.
- Twirrim 2y agoThey're not on AWS. They use on-prem and are migrating to Google and Oracle clouds. https://www.forbes.com/sites/danielnewman/2023/02/21/uber-goes-big-on-cloud-with-google-and-oracle-as-cloud-architecture-debate-continues/ https://www.forbes.com/sites/danielnewman/2023/02/21/uber-go...
- remon 2y agoIt's sort of funny how can you immediately tell it's LLM sanitized/rewritten.
- deleted 2y ago[deleted]
- 1f60c 2y agoI got that feeling as well. In addition, I suspect it was originally written for an internal audience and adapted for the 'blog because the references to SLOs and SLAs don't really make sense in the context of external Uber customers.
- cheema33 2y ago> it's LLM sanitized/rewritten LLM is the new spellchecker. Soon we'll we will wonder why some people don't use it to sanity check blog posts or any other writing. And let's be honest, some writings would greatly benefit from a sanity check.
- bronzekaiser 2y agoScroll to the bottom and look at the authors Its immediately obvious
- karthikmurkonda 2y agoI don't get it. Why is it so?
- lawrjone 2y agoYeah I found this really off putting: it’s not possible for you to have several goals that are all ‘paramount’, and the word ‘seamless’ adds nothing in every place it appears! I wish it didn’t turn me off the content as much as it does but it’s very jarring.
- aprilthird2021 2y agoLet's delve into why you think that
- jeffbee 2y agoFile under "things you will never need to do if you use cloud services".
- mannyv 2y agoThat's not true. The RDS 5.7 instances are EOL so you have to upgrade them at some point. At least in RDS, that will be a one-way upgrade ie: no rollback will be possible. That said, you can upgrade one instance at a time in your cluster for a no-downtime rollout.
- jeffbee 2y agoHosted MySQL is not what I meant. That just means you're paying more to have all the same problems. The kind of cloud service I am alluding to is cloud spanner, cloud bigtable, dynamodb.
- martinsnow 2y agoNah random outages because the RDS instance you were on decided to faceplant itself, or the weird memory to bandwidth scaling AWS has chosen will make you pull your hair out on a high traffic day. It's just different problems.
- jeffbee 2y agoThe company in the article is doing < 200qps per node. Unless they are returning a feature-length video file from every query, they are nowhere near any hardware resource limits.
- paxys 2y agoAt Uber's scale they are a cloud service.
- anitil 2y agoI wonder why they did a large version jump in one shot (v5.7->v8) and didn't do incremental upgrades (v5.7 -> 6.x etc)? I wonder because the promotion of the secondary v8 node to primary is a breaking change in this path, whereas in an incremental upgrade it might not have been. But I also understand at this sort of scale things might be as easy as that.