4 ms·
Maybe for some use cases but most people on this site are concerned about scaling web sites where scaling up (which is what you suggest) isn't feasible because
by tobi 18y ago
Maybe for some use cases but most people on this site are concerned about scaling web sites where scaling up (which is what you suggest) isn't feasible because of the obvious cap to the approach and the fact that you can't even add your own hardware to cloud/virtual hosting which is quickly becoming the norm.
In general, every problem that can be solved in software should be solved in software.
Postgres can be the second coming of jesus but it's still utterly infeasible for high traffic web applications because of the replication issue. Even the current hacks that add replication will not work because they cannot deal with schema changes without downtime. At Shopify we add an average of 12 columns and 2 tables to the database every month and we had a grand total of 58 minutes of downtime in 2008.
- aditya 18y agoTobi, that's good to know... I'm assuming you're running MySQL then?
- tobi 18y agoWe actually launched Shopify beta on Postgres. One of the last things we setup before launch (literally 2 days before) was replication. This was foolish by us but there was a lot on our plates back in the day. We spent the entire weekend setting up every different replication solution for postgres and in the end I just had to make a judgement call and went to mysql. Luckily we used rails which is pretty agnostic and we have a huge unit test coverage so we could quickly get it ported. We converted the data and launched on MySQL and have been very happy with it ever since. In fact this was a good decision for other reasons as well: the MySQL query cache saved us from having to seriously look at caching for a crucial 3-4 months which we could spend on more important things. On a personal note, i like Postgres much better if it weren't for those issues.
- xzilla 18y agoWow! For less than $5000 and a days work, you could have hired any number of postgres consultants to setup PITR warm standby for you (or Slony if you really thought you needed the read slaves). Instead, you migrated your entire architecture over to another database system. Doesn't sound like a good idea to me, but I have to give you points for chutzpah!
- tobi 18y agoFirst of all, at the time 5k was completely unrealistic. I've worked for nearly 20 months without salary on Shopify before launching it and all my last savings went to pay (very low) salaries to my friends and colleagues who agreed to work on the project because they were passionate about it. But even now, Shopify being a multi million dollar company, I still don't think this would have been well invested money. We had Slony setup and PITR is of no use here either. The point is that you cannot upgrade the schema and data in a migration without downtime in a Postgres setup. Mysql simply replicates alter table statements to it's clients and everything stays in sync. Besides, changing architecture took - as i said - a few hours which ends up even at crazy hourly rates to be no more than 1k so even by direct comparison we saved money. I know that you mean well with your suggestion but it's the same thing i've been hearing from a lot of Postgres supporters, they always argue that we did something wrong.
- xzilla 18y agoYikes... didn't mean to make you so defensive! Let me retort and clarify (hopefully with understanding that this is not meant as criticism) :-) First, I know people doing rails+slony; I think the most common way is by piping the sql to a file and then to slonik, but there are other methods too. Granted, it's not awesome. For a new rails/postgres shop, I'd look at pgpool (provided you really needed replication, which is dubious for a lot of people) Second, this wasn't really about postgres/mysql. If you had told me you went from MySQL->Oracle 2 days before launch, I'd have concerns. Heck, even going from Postgres 8.2 -> 8.3, I'd want to make sure you had good test plans (which it sounds like you did). As a general rule, swapping out your database infrastructure is not something I recommend people take lightly. Yes, rails shops have already decided that the application code is more important than the data, so it's a lot easier to do, but given subtle differences in SQL implementations, people still get bitten by it. Please note, I never said you did anything wrong. You did something risky. Most people don't succeed with risky (which is why it doesn't sound like a good idea to me), but since you did I have to give you credit (again) for making a ballsy call and pulling it off. But I think even you realize that you probably would do things different if you had to do it over again.
- moe 18y agobut it's still utterly infeasible for high traffic web applications because of the replication issue Sorry, but what a nonsense. High traffic sites scale by caching, sharding and non-relational databases, in that order. Replication can, at best, be an intermediate kludge for read-heavy websites that haven't learned about caching (wink) or insist on abusing their RDBMS as a fulltext search engine. It was also commonly used as an incremental backup solution until filesystem snapshots became commonplace.
- aditya 18y agoI think having multiple read slaves (and even write masters) is a perfectly acceptable way of distributing load, and in turn, scaling. To completely ignore that sounds foolish. EDIT: xzilla is right, I wasn't really recommending blindly adding write masters since that wouldn't work, but just that read slaves are not a bad idea. :)
- xzilla 18y agoAnyone who thinks you can scale writes by adding more masters is plain ignorant on the subject. The only way to scale writes is with better hardware or federation of application/data. The first of those is easy, the latter is hard.
- moe 18y agoWell, I have never seen a read-slave setup in a webapp for the purpose of scalability that would've made a lot of sense. It's a sometimes a stopgap measure when money is cheaper than time, but will come back to haunt you later. When your reads are starving in a webapp then you should look into fragment caching and content-generation but not much further. If that doesn't solve your problem then you're just very likely doing something fundamentally wrong and better go ask someone smarter.
- tobi 18y agoReplication has nothing to do with scalability in most cases. Replication is how you reduce downtime! It's the same reason why Raid5 is more then useless in a web server - if a disk dies it has to reboot and has to reconstruct data for hours. You cannot have downtime in a web application. This is of course slightly domain specific. We host several thousand e-commerce stores that are the livelihood of our clients. If we are down then no one can earn money. Shopify being down is a lot worse then twitter being down. In fact we may even see legal action if our downtime is too bad. The reason why you need replication is because you need 3+ 100% accurate and up to the second copies of your database which can take over at a moments notice. One single server, no matter how beefy it is, can never accomplish this. Someone is going to trip over it's power cords eventually ( and if it has two then someone is going to trip over the power cord of the network switch it's connected to ).
- Andys 18y agoWhen I was using pgpool2, it dealt with schema changes seamlessly - you connect to the proxy and issue your table updates, and it all just works.