4 ms·
To make it easier to think about databases, I group them into categories - SQL Databases (Mysql, Postgres, SQL Server, SQLite) - JSON document stores (MongoDB
by twunde 4y ago
To make it easier to think about databases, I group them into categories
- SQL Databases (Mysql, Postgres, SQL Server, SQLite)
- JSON document stores (MongoDB, etc)
- In-memory key-value stores (redis, memcache)
- key-value stores (cassandra, etc)
- Graph DBs (neo4j, dgraph)
- Search DBs (elasticsearch, Lucene)
- Timeseries DBs (timescale)
I consider the SQL databases and the JSON document stores to be the general default databases that my applications will use 90%+ of the time. SQL databases require defining the schema ahead of time, while with document stores I can just dump data into the database and then figure out what's needed afterwards. At this point, I usually use postgres as my go-to database since it's open-source, well-supported and has some additional features like the ability to use json, but if you're more familiar with a different SQL db use that instead. Most of the other database categories are specialized for specific use-cases (in-memory key-values are great for queues and caching, time-series dbs are great for metrics)
EDIT:
https://db-engines.com/en/ranking https://db-engines.com/en/ranking currently shows 381 unique databases, which is probably a bit of an undercount but not by much.
I realized I didn't answer the `Why there are so many databases` part of your question. Essentially, many companies outgrow the general-purpose databases and need databases optimized to their read-write level and data structures. Analytics tend to use columnar-databases once their data sets are in the TB+ range so that their queries return quickly.
- sargstuff 4y agoProject metrics & usage / design doumentation should be able to help with selecting relevant database grouping(s). aka post data collection doing timeseries analystics; at home book library database so, no need for remote web access; doing local gis map database; database for motion sensor to record time / date motion detected. starter ideas from searching on "database selection metrics" after:2019-12-31 short reads: https://medium.com/wix-engineering/how-to-choose-the-right-database-for-your-service-97b1670c5632 https://www.pingcap.com/blog/how-to-efficiently-choose-the-right-database-for-your-applications/ https://dzone.com/articles/introducing-database-selection full discussion: https://sunnykichloo.wordpress.com/2020/09/21/database-selection-criteria/ https://www.mongodb.com/blog/post/introducing-database-selection-matrix
- sargstuff 4y agoInternet search engines are 'the' widely used, general purpose database, at least for things that can be shared publicly. Obviously, doesn't work to well with embedded devices without an internet connection and/or data not in some electronic format aka traditional camera picture / computer file
- gregjor 4y agoPlain files will work for those collections at the kind of scale implied. Don’t overthink it.
- sargstuff 4y agohence the comment on project metrics / usage / design. aka aggregating a few terabytes of motion data may not work well in a single text file for CERN collider motion data. ;)
- gregjor 4y agoRight... you previously listed a home library catalog, local GIS data, motion capture for what, a home Nest device? You should have mentioned you need a solution for CERN collider data or that you have a collider at home for more relevant answers.
- sargstuff 4y agook, should have prefaced with, miscellaneous contrived use case examples, where project metrics/usage/design intormation would provide insites about the contrived use intent / framework.