4 ms·
Postgres isn't designed from the ground up for high availability or scalability. Most of the in-built functionality for setting up HA wasn't introduced until 9
by coderzach 11y ago
Postgres isn't designed from the ground up for high availability or scalability. Most of the in-built functionality for setting up HA wasn't introduced until 9.0, and the process of setting up HA postgres with automatic failover requires third-party tools. As for scalability, postgres can only scale vertically.
If you need a solution that can scale writes to more than than one node, or a solution that has first party support for HA with automatic failover, you shouldn't be using postgres.
As an aside, I find the dogmatic "just use postgres for everything" choir just as bad as the marketing BS associated with noSQL databases.
- snuxoll 11y agoUnless you are using synchronous replication (locks up transactions if there is an issue with replication) there's no safe way to do automatic failover with an RDBMS without using shared storage (expensive, pain in the butt). Corosync and pacemaker make setting up automatic failover with shared storage relatively painless. I don't do it this way personally because shared storage is a giant pain.
- empthought 11y ago"High availability" and "high availability with automatic failover" are not the same thing. (The tricky bit is actually the un-failing-over part.) I assure you that people were setting up high availability systems with PostgreSQL 7.x in 2004 or so. Very few organizations need to scale their database in any other direction than vertically. If it works for Stack Overflow, it's certainly likely to work for your application.
- takeda 11y agoIt's also lack of understanding the problem you're trying to solve. For example at my previous job we had a dedicated data store for performing lookups such as ip -> zip code and latitude/longitude -> zip code. The company decided to use Oracle Coherence and store all data in memory, because it'll be fast. To store all of that information they needed 16 m3.medium machines. Last year they had a great success optimizing it, because they managed to replace 16 x m3.medium machines to just 3 x c3.2xlarge machines running MongoDB (the data was ~12GB). I did a POC and put the data in PostgreSQL with proper columns and indices (I just needed to install ip4range and PostGIS), the whole data fit in 600MB! the queries took under at most 2ms on cold cache but generally were under miliseconds, because all the data fit in RAM.
- threeseed 11y agoIt seems more of a problem with the people than technologies. Why would you need 16 m3.medium (60GB RAM total) to store 600MB of data ? I've used Oracle Coherence and other grid technologies and something doesn't sound right here. It is just a couple of distributed Java HashMaps we are taking about here. Likewise if 600MB of data is expanding to 12GB in MongoDB then something is very, very wrong with the design of your schema.
- empthought 11y ago> It seems more of a problem with the people than technologies. Isn't that what "lack of understanding of the problem you're trying to solve" means? Though in this case, if the problem had "IP4 ranges" and "geographic data and computation functions" in its scope, then MongoDB is quite inadequate compared to PostgreSQL.
- takeda 11y agoYes problem were people. Too much politics and that's why I left. Why more data? In case of IP geolocation neither of that technology understood IPs not to mention being able to create a proper index for ranges. So in case of Mongo, to get a good performance they decided to generate every possible IPv4 address and map it to zip code. To increase efficiency they stored every IP as a 64 bit integer. In Coherence they did the same thing, but I guess less efficiently (did not look how it was done, since at the time coherence was in the process of being eliminated) I'm guessing maybe they stored is as a string? Also note that Coherence is a distributed cache that supposed to withstand couple nodes going down, so a lot of data was duplicated.
- threeseed 11y ago> Very few organizations need to scale their database in any other direction than vertically Almost every major organisation and plenty of startups/SMEs are investing big into analytics programs. And whilst they don't have massive data sets they are expecting real time performance so being able to scale horizontally is important. If you could vertically scale I/O that would be one thing but you can't.
- snuxoll 11y agoPL/Proxy?
- empthought 11y agoThat's a bunch of read-only replicas though; RDBMSes can do that. Or if it's really big data, then we are talking about Redshift, Greenplum, or Teradata. There's no need for something like CouchDB.