8 ms·
Thoughts on Uber’s List of Postgres Limitations
- bechampion 10y ago"Error establishing a database connection" ^ Running on pg too?
- ZeKK14 10y ago"Error establishing a database connection" Genius !
- Sevrene 10y agogoogle cache: https://webcache.googleusercontent.com/search?q=cache:blog.2ndquadrant.com/thoughts-on-ubers-list-of-postgres-limitations/&num=1&strip=1&vwsrc=0 https://webcache.googleusercontent.com/search?q=cache:blog.2...
- andy_ppp 10y agoMight be worth someone writing "a panicky guide to installing varnish" for such database issues. Pretty embarrassing though!
- grabcocque 10y ago'Error establishing database connection' How very meta.
- dozzie 10y agoHNHOD: Hackernews Hug of Death.
- taspeotis 10y agoFortunately they can call themselves! WHEN IT'S CRITICAL, YOU CAN COUNT ON US 2ndQuadrant provides full 24/7 problem resolution technical support for production systems. If PostgreSQL breaks, we'll get you back up quickly. [1] https://2ndquadrant.com/en/support/support-postgresql/ https://2ndquadrant.com/en/support/support-postgresql/
- simon2Q 10y agoWhich is exactly what happened and we fixed it fairly quickly, within the SLA we offer to customers. So I'm happy. I accept all of the humour on that point with a grin myself, though I must say its a nice problem to have. Thanks to everybody for reading and commenting. 2ndQuadrant is a large enough company that we have CTOs who write blogs and design stuff, we have other people who run blog websites and a variety of infrastructure, but mainly we have many dev and support staff helping customers.
- awalGarg 10y agoThe site is powered by Wordpress, which, to the best of my knowledge, can work only with MySQL. Now that is very meta.
- epimetheus 10y agoIt can work with MariaDB as well, though it's just a Fork of MySQL so hard to say if that counts.
- simon2Q 10y agoI believe the blog site does run MySQL. We eat our own dogfood wherever possible, especially on important services... we hadn't regarded the blog site in the same category until now.
- Illniyar 10y agoI think its worth mentioning again that what Uber ended up using has no resemblence to an RDBMS (single table, manual indexes). So regardless of whether their complaints are justified or not, it should not be taken as an endorsment of mySQL over postgres, but rather of an endorsement of NoSQL over RDBMS . Which is really just what every company at these scales do (except for google and f5 if whitepapers are to be considered).
- leftnode 10y agoCan you expand on what you mean by a single table with manual indexes? How does a frequently used table with indexes not resemble an RDBMS?
- Illniyar 10y agothe secondary indexes are saved in a different table and (I can only assume) the application layer (or one layer above mySql) is responsible for keeping that table up to date. If you use only 1 table, then it's by no means "relational", so I don't see why you'll need a database system designed from start to finish to support a relational model.
- pas 10y agoThe "relationalness" is/was not really important here. It's all about MVCC and how storage engines handle it. Postgres is lacking in these scenarios, whereas a particular fine tuned version (or fork) of InnoDB (or MyRocks or whatever they end up chosing) handles this better. See Facebook's "mysql-5.6" branch, that has hundreds of patches piled on to support especially these taxing workloads.
- Illniyar 10y agoWell, for that matter, as I understood it, they also have no transactions nor atomicity and are basically eventually consistant , though these I'm simply inferring from their posts, it was not stated as far as I can remember. So MVCC also has almost zero bearing on what they are doing.
- Dowwie 10y agoDoes anyone know the back story to Uber - why didn't it try to improve Postgres rather than move on to feed on another host?
- tehbeard 10y agoMost likely they don't have any engineers with the skills necessary to build/improve a RDBMS. (note I said build, not use, different skillsets between driving a car and re-configuring the engine to run on seed oil)
- deleted 10y ago[deleted]
- okket 10y agoWatch https://vimeo.com/145842299 https://vimeo.com/145842299
- merb 10y agoactually I watched the video and the errors happend here are more cultural errors. I guess the switch is more like a "we failed, so we start new". and new means with something different. also their replication strategy looked like a joke and not enough automation. hopefully they never use galera, else their bad engeering practices could actually suffer in huge data losses. I've once run galera-cluster and it pretty easy came to data losses especially after short network splits which occured randomly at the network.
- pwg 10y agoIn one of the other HN threads on the issue, there was mention that the switch occurred shortly after a change in leadership position of individuals with enough 'power' to impart technology changes by decree. So there is an aspect that there was some, if not much, political opinion involved as there was any technical analysis involved in the change. Given the fact that the change which occurred is an "apples to oranges comparison" change (PostgreSQL and an SQL normalized DB table structure to MySQL and a noSQL style single table key-value store) then there is some credence one can put towards the rumor that some (or most) of the change may have come about from a PHB [1] saying "do this, this way". [1] PHB (Pointy Haired Boss, Dilbert cartoon reference)
- f055 10y agoI came here to write a snarky comment, but now I can write two ;) first: if you think any particular db platform is clearly a winner in "db wars", you are naive. there are so many factors involved in configuring the db, the backend, the frontend etc. that you can always find a case where: the supposedly winning db is failing, or the supposedly worse db is performing perfectly fine. and from my experience, you should always use the platform/framework/language that is best for the current project, not the one you madly love. clearly postgres wasn't working for uber. that does not mean it will not work for your project. i have a recent experience where a binary file of a programming object works much much faster than mysql and solves several other problems. would i say "use binary files instead of rdbms"? of course not. but in this one case it does wonders. the "tech-vs-tech" wars need to end, they are pointless. second: if you cannot setup your blog to withstand an HN spike then maybe you don't have as much real world experience with scalability (albeit simple) as you might think (hint: static page cache behind cdn will make you almost bulletproof - also, with for example Azure, that's dead cheap).
- dkersten 10y agore: your second point, this isn't necessarily anything to do with database scalability or tuning.
- f055 10y agoexactly my point. sometimes the solution has nothing to do with the perceived problem.
- eyan 10y agoI came here to thank the Postgres developers for their hard work. Also their grace under such trying times.
- majewsky 10y ago> static page cache behind cdn I wouldn't want my page behind a CDN. CDNs make users much more trackable across sites. My point isn't that CDNs are bad for everyone. My point is, once more, that most questions are not as simple as they may appear.
- harshal 10y agoPotentially the most useful part of this post to me was this part > 2ndQuadrant is working on highly efficient upgrades from earlier major releases, starting with 9.1 → 9.5/9.6. I hadn't heard of that before. Anybody know more about this? I'm currently babysitting a 9.1 deployment which we desperately want to get upgraded. The amount of downtime this can tolerate a very limited and I was currently tasked with coming up with a plan. Its going to get hairy. If such a tool is really on its way, I could make a case for holding off on the upgrade for a few more months and save quite a bit of work.
- pilif 10y agoyou can use pg_upgrade with -k - it will complete within seconds. Afterwards, things will be slow until a complete analyze updates the statistics, but the update itself can be done in seconds. I have updated ~2TB of database from 9.0 all the way to 9.5 over the years.
- bigato 10y agoThe problem with this is that if anything fails, you can potentially corrupt your data and have no backup plan. To make that option safe, you would have to copy your data directory first, and you need to be offline for that. So you have to add the time it takes to make that copy.
- pilif 10y agoThis is why I ensure that the slaves are up to date, then disconnect them, pg_upgrade the master and resync the slaves (which is required anyways). If something goes wrong, I would fail over to the slave. Also: You don't need to be offline to copy the data directory. Check `pg_start_backup` or `pg_basebackup` (which calls the former)
- bigato 10y agoThat requires the master and slaves to run different versions for a while. And that is not possible with stock postgresql, is it? Regarding your second point, I meant copying the data directory as in a 'cp' command. Or rsync if you will. The functions you mentioned are only useful when doing a dump, isn't it? And recovering from a upgrade problem using a dump is way slower than just starting the previous version in the backup data directory.
- ungzd 10y agoIt seems that after lots of hyped NoSQL systems companies are still using MySQL or Postgresql as simple storage backend with lots of custom crutches over it, like it's 2007. So all these cassandaras and riaks failed expectations?
- threeseed 10y agoThis is one of the more ridiculous posts I've seen on HN. "Companies" are taking a whole range of approaches to storing data. With a combination of NoSQL, SQL and Filesystem e.g. HDFS and everything else in between. Cassandra in particular is killing it right now under the stewardship of Datastax which is why they've grown from 1 person up to 400+ employees.
- snaky 10y agoCompanies, outside the startup scene, are still using Oracle and MS SQL Server, and for the most part are pretty happy with it.
- 3manuek 10y agoI do think that you need to learn how to use in the best way the technologies you have chosen or that are present in your current setup. No matter if it is a MySQL, Postgres or any other DB, it requires as a part of a job, learn how to use at it's best. The points on the article are good, however it's true that Postgres had problems with scalability not so long ago. That's changing, however I think that other data stores have addressed the problem of availability and scalability earlier and gained maturity during the last years. Also, there is something that caused some noise to me: This point is correct; PostgreSQL indexes currently use a direct pointer between the index entry and the heap tuple version. InnoDB secondary indexes are “indirect indexes” in that they do not refer to the heap tuple version directly, they contain the value of the Primary Key (PK) of the tuple. That's true, but the article doesn't make explicit that the PK on InnoDB is a clustered index and, that there are other optimizations like adaptive hashing to make read queries faster.
- evanelias 10y agoAgreed 100%, and the author also failed to mention several other advantages of having a clustered PK and indirect secondary indexes. A few off the top of my head: reads in PK order are faster due to lack of indirection; clustered index takes up less space due to lack of storing pointers to tuples; secondary indexes will be covering (no need for PK lookup at all) if the query only uses columns in the PK and secondary index. It is interesting/ironic to see the article complain "those limitations were actually true in the distant past of 5-10 years ago, so that leaves us with the impression of comparing MySQL as it is now with PostgreSQL as it was a decade ago." In the MySQL world, we very very frequently see the opposite -- Postgres fans bashing MySQL for things that haven't been true in 10-15 years, as well as things that simply have never been true. It certainly is frustrating, just like what the author is experiencing! Having a favorite/preferred database is fine, but I don't understand all the extreme views -- why do so few of these articles take the view that Postgres is a better fit for some workloads, and MySQL/InnoDB is a better fit some other workloads? Or even just an acknowledgement that the authors of these articles rarely, if ever, have a comparable amount of expertise in both databases -- which would be necessary to make a fair comparison. Yes, Uber's original article clearly shares this same problem, but at least they seem to acknowledge it more clearly than the author of this response article. Take the section on replication comparison, for example: the author is describing logical replication support in Postgres even though it's currently a third-party addon. Cool, but MySQL has all sorts of third-party replication systems too. Alibaba has implemented physical replication in MySQL. And meanwhile even in MySQL core, there are two different types of logical replication -- there's no restriction to only use statement-based logical replication as this article implies.
- asah 10y agoTo me, the real news is that Uber ($50B company) didn't bother to engage the postgres community before migrating - they'd have jumped to support Uber.
- rch 10y agoI don't get the sense that Uber has had much stability in technical leadership over its lifetime. Big moves like this can be as much cultural as technical.
- snaky 10y agoWhat would be better to have in resume, "we were using PostgreSQL and after some tuning it worked just fine" or "we designed and implemented scalable modern BigData realtime OLTP solution"?
- known 10y agoDoesn't Postgres use mmap files internally?
- anarazel 10y agoNo. It uses mmap(MAP_ANONYMOUS) for most of its shared memory, that's pretty much all the use of mmap.
- raisingqs 10y agowould simply coughing up the money for an expensive oracle/sql server solution have worked in this case?
- raisingqs 10y agowould simply coughing up the money for an expensive oracle/sql server solution have worked in this case?