7 ms·
Why Uber Engineering Switched from Postgres to MySQL (2016)
- thirduncle 8y agoFrom 2016 and definitely a repost.
- peterwwillis 8y agoHere are some of the old comment threads: https://news.ycombinator.com/item?id=12166585 https://news.ycombinator.com/item?id=12166585 https://news.ycombinator.com/item?id=12179222 https://news.ycombinator.com/item?id=12179222
- sidcool 8y agoThis is a 2016 article, can someone add the timestamp?
- dotdi 8y agoMaybe it's just me but I find this article has the following TL;DR: we are going against recommendations and common practices and postgres doesn't work well.
- hestefisk 8y agoIf what they wanted was a Schemaless, sharded data store, then perhaps Postgres isn’t the best or most fair comparison. Also the claim of issues with table corruption seems strange / unbelievable— I have used Postgres as my daily driver for data (at scale)!for 12 years now and yet to see it. But again, I haven’t got same scale requirement as Uber.
- deleted 8y ago[deleted]
- sitharus 8y agoUse any database enough and you'll find weird table corruption bugs. I've seen them in MySQL and MSSQL in my professional career. Fortunately in both cases it was caught before there was an unrecoverable issue, but it happens. I've read the article before and I do agree, Postgres wasn't really the tool for what they're doing.
- pilif 8y ago> Use any database enough and you'll find weird table corruption bugs I'm not convinced. The one feature I really need a database to have is to guarantee that whatever data I write to it, I will be able to read back unaltered. I'm a heavy Postgres user since 2001 and I have never seen table corruption, even in the face of defective hardware. The only time something remotely comparable has happened was index corruption on a replica due to a bug in 9.2.0 which was very quickly fixed and as it was just a corrupted index, very easily restored (REINDEX). If a database I trust with production data was corrupting tables, I would move off that system in a heartbeat (which is incidentally why I stopped using MySQL back then)
- debaserab2 8y agoThe missing context here is the scale. What size of data and transaction volume have you been working with?
- pilif 8y agoBy now the data is about 6.2 TB with overall 20k transactions per second, so I would say not small any more, but also certainly not Uber scale. Still. I have had MySQL corrupt tables with way less load and data (in the less than 10GB size range)
- jjirsa 8y ago6.2TB is single machine scale The company here is doing petabytes of data and millions of ops per second. That’s scale. We should be clear about what “at scale” means, and 6.2TB is not.
- pilif 8y agoNo matter the scale. Table corruption is never acceptable. If table corruption is acceptable, you might as well just serve data from /dev/urandom and store user data to /dev/null.
- sureaboutthis 8y agoIf Schemaless worked on PG, would this article have seen the light of day?
- hestefisk 8y agoThere are solid alternatives. pg_json with Postgrest on top. I should not have capitalised Schemaless as I was more referring to the notion of a schema-less database.
- akanet 8y agoI think Evan's identified a ton of problems with Postgres (at the time) that extend far beyond just the query/data pattern. Many of these problems were directly acknowledged by the PG team, which made for interesting reading: https://www.postgresql.org/message-id/flat/579795DF.10502%40commandprompt.com#579795DF.10502@commandprompt.com https://www.postgresql.org/message-id/flat/579795DF.10502%40....
- muddyrivers 8y agoThe table corruption bugs were triggered when the load was very high, as far as I know. It happened twice to me. In both cases, the corruption couldn't be corrected. The worse part was the corruption was propagated into replica, which brought up much more serious issue with Postgres's replication model. So the only solution was to install the latest good backup.
- iKSv2 8y agoAlthough this is an old article (maybe even posted here earlier), curious to know how's does pg10 or upcoming pg11 serve / address those issues mentioned....
- jasonjayr 8y agohttps://news.ycombinator.com/item?id=12166585 https://news.ycombinator.com/item?id=12166585 Previous discussion from 2016 ...
- tananaev 8y agoA couple of years ago we did a load testing comparison between MySQL and Postgres for our Java-based project. MySQL throughput was 5x better. I'm not sure if it was JDBC driver issue or just database configuration wasn't perfect, but MySQL worked much better for us.
- NickNameNick 8y agoThe default PostgreSQL configuration is extremely light-weight. Allowing it to use (much) more memory is a good start. https://wiki.postgresql.org/wiki/Tuning_Your_PostgreSQL_Server https://wiki.postgresql.org/wiki/Tuning_Your_PostgreSQL_Serv...
- setquk 8y agoWe've got some non production facing stuff being hammered with the default postgres config. It's actually surprisingly good out of the box. We run on the basis that we will tune it if we need to with that and we haven't had to yet.
- pritambarhate 8y agoI don't think there is any issue in PgSQL JDBC drivers. In fact in the latest round of TechEmpower benchmarks[0], if you filter for Java, you will see all top positions are held by PgSQL based stacks for Single Query, Multiple Query and Updates tests. 0: https://www.techempower.com/benchmarks/#section=data-r16&hw=ph&test=update&l=hra0e7 https://www.techempower.com/benchmarks/#section=data-r16&hw=...
- zaxomi 8y agoFrom the article: > MySQL’s replication architecture means that if bugs do cause table corruption, the problem is unlikely to cause a catastrophic failure. Replication happens at the logical layer, so an operation like rebalancing a B-tree can never cause an index to become corrupted. A typical MySQL replication issue is the case of a statement being skipped (or, less frequently, applied twice). This may cause data to be missing or invalid, but it won’t cause a database outage. So, data might be invalid or missing, but it will not cause a database outage...
- sandGorgon 8y agothis is a huge accusation against Postgres and at Uber scale, I would assume some homework has been done here. Does anyone know if this is relatively true - within the strict constraints of comparing mysql and postgres ? Is mysql really prone to lesser corruption than postgres.
- dijit 8y agoMySQL is prone to /more/ corruption, not less. It's gotten better in recent versions, but implicit type conversion and truncation of data is a common "feature" of MySQL. 0000-00-00 isn't a valid date; but in MySQL it is a "null" data on a column with a NOT NULL constraint. PostgreSQL is more anal about what you put into it, if you try to insert into a row with no value for a NOT NULL constrained column it will fail the transaction. Once data is in the system mysql will tend to favour binary replicated statement level queries; where postgresql's built-in replication is block level. This means that if you delete a row on a mysql replica the replication will continue and you will only notice that you're out of sync when you try to alter the (now missing) row from the master; then replication will fail.. It's impossible to even modify data on a postgresql replica. My understanding is that mysql will also truncate data in a column if the column is altered to where a value would no longer fit (IE; from varchar(20) to varchar(8)) where postgresql will fail and abort the transaction. (This is my experience from using PGSQL and MySQL with varying degrees of scale for 15 years). PostgreSQL has/had problems to be sure, the auto vacuum was an issue on the 8.0->8.2 line; and they're slow to adopt features (like baked-in replication and upsert) but I treat this as a feature itself; better to have a working feature that is cleanly engineered than a feature which is half-baked.
- samgranieri 8y agoThis is old news and a repost. pass...
- Operyl 8y agoWhile it is old to you, HN prefers to add a year to the title so that others might happen upon it if they missed it in the past. Not everyone is able to watch HN 24/7, and new discourse might have come up since the last time this was posted.
- watt 8y agoPostgres however has moved quite a way since 2016 when the post was written, so it is possibly even factually incorrect at 2018. Postgres switched to much more rapid release scheme, and is making a lot of ergonomic and performance improvements since they got the "correctness" part down (such as auto vacuum, etc), which only came together in last couple years.
- Operyl 8y agoI feel like the 2016 tag is extremely useful in letting users figure this out, though. And the new discourse that flows from this (read some of the other comments in here, they're explaining that Postgres took a hard look at themselves and adapted).
- owaislone 8y agoExcept you stopped to comment :)
- staticelf 8y agoSince I started programming, I have always used MySQL because honestly I think it is the best DB I have used because it has great software support and usability. I have tried Postgres and while I have no issues with it, I usually go with MySQL if I have a choice because I see no real reason to switch out something I know and like. Is there any real reason for a person like me to use postgres or another alternative instead of MySQL?
- BozeWolf 8y agoMany reasons, full text indexing is a good one for example. Although mysql claims to have support for it, it is not as far as postgres is (with tsvector). It has many more extensions. It depends on what frameworks you use, but django is mostly optimized for postgres. The MySQL ecosystem is a bit weird due to having multiple "mysql"-ish versions. I did a small project where mariadb and mysql where not interchangeable. Some stuff was causing problems, cannot remember exactly anymore. Edit: EXPLAIN ANALYZE in postgres I love it.