4 ms·
Not sure about all of them but different databases have different use cases. It's not necessarily Postgres vs. MySQL where they are competing databases for the
by dmlittle 6y ago
Not sure about all of them but different databases have different use cases. It's not necessarily Postgres vs. MySQL where they are competing databases for the same use case. Snowflake, for example, is a data warehouse that (I believe) stores data in a columnar format. This makes it much better at performing aggregate queries (how many shows were watched on netflix today?) but a lot slower for specific queries (where in this episode is gavinray currently at?).
- gavinray 6y agoHive is a Data Warehouse that runs on top of Hadoop. Snowflake is a Data Warehouse/Lake, but it's also it's own custom SQL DB. Cassandra is a NoSQL DB RDS is running (some standard relational DB) Apache Druid is a columnar analytics DB centered around realtime uses. It has it's own query language and delegates to Apache Calcite for specific DB/datasource drivers. Can integrate with Kafka/Hadoop. Presto is (to my understanding) like a meta-DB that can query multiple databases. Similar to an integrated Apache Calcite or Google's ZetaSQL. There are a LOT of overlapping concerns here, which is why the confusion. Essentially they have 2-3 different products in several categories targeting generally the same usecase.
- texasbigdata 6y agoIs there a good resource for this comparison plus the many streaming/ETL permutations. It's a bit confusing even just what the various Apache products do (beam, kylin, etc, etc).
- nemothekid 6y agoAFAICT, Hive/Snowflake are the only overlapping concerns here, and I can see why they would have both (Snowflake may be the better product, but at the end of the day Netflix owns their Hive install). Everything else solves different problems at Netflix scale. I'm spitballing here but Cassandra could be used for metadata serving (high throughtput, embarassingly parallel reads with high uptime), RDS for their billing system (transactions, ACID, etc), Druid for realtime OLAP and Presto as an interface to Hive/Snowflake. A smaller company wouldn't need this level of complexity (if you aren't large, you could probably serve your metadata from MySQL, and just use Snowflake as your OLAP engine).
- ramraj07 6y agoHuh, even MySQL and postgres are not apples to apples - we are just now launching a service in MySQL instead of our standard postgres because MySQL has features that postgres doesn't (better connection pooling, potential data partitioning)
- dmlittle 6y agoI would probably say MySQL and Postgres is apples to apples just different types of apples (granny vs delicious). Their differences are not as big as a regular relation DB vs. a columnar DB.
- gshulegaard 6y agoI am sure MySQL is a great fit for your use case. For others reading your comment though, I did want to list some things I have used with Postgres that relate to connection pooling and data partitioning: * PGBouncer for connection pooling/sharing. * Postgres Table Inheritance for table partitioning (https://www.postgresql.org/docs/12/ddl-inherit.html https://www.postgresql.org/docs/12/ddl-inherit.html) * PGPartman for automating the creation of partitions (https://github.com/pgpartman/pg_partman https://github.com/pgpartman/pg_partman) * Citus for low-barrier data sharding (it's a Postgres Extension like PostGIS) (https://www.citusdata.com/ https://www.citusdata.com/)