19 ms·
Why NoSQL
- einpoklum 5y agoThis article regards write-intensive usage scenarios (OLTP) rather than analytics-focused usage scenarios (OLAP).
- nimchimpsky 5y agoI'll be honest they lost me at "client side, offline first, JavaScript database". And then they lost me again when they didn't format hrefs different to normal text so I couldn't copy the text.
- zz865 5y agoI thought nosql would make my life easier but it doesn't. Use RDMS with an upfront schema really is simpler. Yeah changes can be awkward but its worth it.
- eru 5y agoThe schema is basically just static typing for your data. So the same pros and cons as for static vs dynamic typing in programming languages apply. (As an interesting aside, I used to work with a system that had relations as in-memory datastructures. They were very pleasant to work with compared to eg Python dicts, because no single one key was privileged, like you have to do with a dict.)
- invalidname 5y agoNoSQL are super easy to start with. Then you have a pile of heaping unstructured data that's hard to query, hard to report and gain insight on. Every data task becomes a programming problem which needs a team to work on to get something that might be stale. With SQL there's a lot of work. But you have a DB that everyone knows. 3rd party tools work. Developer on-boarding is easy. Your boss asks you how many users used feature X in the past month you can actually tell him right away... Performance is actually good and consistent. Your data is also consistent and respects acid principles. The whole write performance of NoSQL DBs is a bit of a crock that doesn't stand the test of scaling. With proper cashing SQL is as performant when done right and the profiling tools are better. Screw no-SQL. They are just Object Oriented DBs (which were a disaster in the 90s) in new clothes.
- YZF 5y agoWhat's your SQL db of choice where you can scale it up seamlessly beyond a single instance? EDIT: to add ... and with HA/DR and where you keep all the good guarantees, queries etc.?
- FridgeSeal 5y agoIf the data couldn’t be fully sharded (I.e. independent db’s and replicas), I’d personally choose Cockroach DB for this, I admittedly haven’t had the pleasure of using Cockroach, but I’ve seen alternative approaches and knowing what I wouldn’t do again…
- qaq 5y agoSeeing most people go for the cloud what percentage of projects will outgrow say Aurora? and beyond Aurora there is a ton of options like CockroachDB, TiDB, Yugobyte etc.
- belk 5y agoCockroachDB, I heard the dev experience of spinning up nodes can be dumbed down to 1 or 2 docker commands, and it's Postgres compatible, continues to have global transactions, although you wont get the same transaction performance as something GPS/clock hardware synced like Spanner, it's still a good self hosted solution. If you're interested, CRDB writes a lot about how they do their transactions on their blog, it's good to know what guarantees you actually get with your DB's transactions.
- nine_k 5y agoFor what reason do you need to scale it beyond a single instance? If all you need is to alleviate the high read load (many selects), nearly every SQL database (even SQLite!) has a free and supported way to create read replicas, usually out of the box. If you need distributed transactional updates, MySQL / MariaDB has Galera (GPL), and Postgres has Citus (proprietary with a few AGPL parts). And, of course, there is CockroachDB that prioritizes reliability over speed, but is truly distributed out of the box.
- 5y ago
- endisneigh 5y agoMy dream is a database where I just add nodes to physical machines by running a docker container and I query using whatever mechanism. From there it would just auto heal, scale and do all the ops by itself. CouchDB is sort of like this in theory (definitely not in practice).
- orangepurple 5y agoThe commercial version of this is Snowflake, except that it's completely seamless, scales invisibly, and its fast as nuts. At least 100-1000x faster on common workloads compared to a mid 2000s IBM SQL DB solution for the same data at a really large enterprise.
- beamatronic 5y agoCouchbase does that
- inopinatus 5y ago> A new client downloads all current documents and each time a document changes, that document is downloaded again Those who do not remember Lotus Notes are doomed to reinvent it.
- slantyyz 5y agoLotus Notes, if you look past it having one of the ugliest UIs ever made, was a pretty neat tool. When I look at a tool like Notion, I see it as what Notes should have become. I worked in a Notes/Domino shop pre-2000, and after some "geniuses" in the company read a Gartner report that SAAS was going to be a thing within a couple of years, my company blew a lot of resources to hack together a multi-tenant project collaboration app using Domino that was delivered to web browsers. Of course, it didn't go anywhere, because it took many more years before SAAS eventually became mainstream. It's funny looking back at what we made around 98-99, because JS on the browser was far less mature than what it is today.
- irrational 5y agoMaybe it is because I've only ever worked on web applications so I can't imagine the use cases, but when would you ever want your database to be replicated to a client's computer?
- er4hn 5y agoOne of the use cases for RxDb is "local first" data. Rather than spend time doing roundtrips to a remote server for each piece of data, store the database locally on the client and keep the local/remote in sync.
- other_herbert 5y agoIndexedDB for the client, web workers to check that things are in sync and events for when things aren’t.. best of both worlds
- tabtab 5y agoTesting and prototyping.
- irrational 5y agoI recently learned DynamoDB (a non-relational database on AWS). It was fairly easy to learn, basically a one day read through the documentation. It made me grateful that I first learned SQL databases and spent years working with Oracle and Postgres. Now I can easily choose the best tool for each job. But, if I had first learned a non-relational database, I could see myself looking at the daunting task of learning about database schema design, normalization, sql, pl-sql, triggers, indices, keys, constraints, etc. and throwing my hands up with the shear weight of it all. It certainly isn't a 1 day learn like non-relational databases. I'd probably try to force non-relational databases to fit every use case, even where they aren't the best fit.
- er4hn 5y agoI feel like I'm missing something in this article. Several of the points just don't make any sense to me. * Their point about how you can just rewrite relational queries manually and not lose any performance does not seem true the moment that you need to join tables on a condition. * Their section on reliable replications can be resolved by keeping a log of the changes made to the database rather than a log of the queries. I don't understand how their NoSql model of "Download the latest doc during a change" is any better. It basically becomes "last writer wins" in a distributed system. * The downgrading to NoSql argument makes little sense to me as well. If your frontend is sqlite then you can have Postgres, or MariaDb in your backend. You need to account for having different queries on the front and the back, but the backend is presumably also operating under very different constraints than the frontend. * I'd actually make an argument that NoSql databases are a type of relational database heavily optimized towards not needing to do joins. The corollary to this is that over time you may find that, oops, you do want to do joins, and now you have to pay a tax on not being set up to do so in the first place. * Also, is it just me or is the Hybrid Logical Clock they mention a type of Lamport Clock where the linked article doesn't ever state that?
- quixoticaxolotl 5y agoHybrid logical clocks are distinct from Lamport clocks in that they include a notion of local timestamp while also preserving partial order. They are well-established and used widely outside RxDB - see e.g. https://sergeiturukin.com/2017/06/26/hybrid-logical-clocks.html https://sergeiturukin.com/2017/06/26/hybrid-logical-clocks.h...
- EdwardDiego 5y ago> Their section on reliable replications can be resolved by keeping a log of the changes made to the database rather than a log of the queries I do love Debezium for easily streaming changelogs to Kafka.
- lmm 5y ago> Their point about how you can just rewrite relational queries manually and not lose any performance does not seem true the moment that you need to join tables on a condition. You can do the same thing a database would do - filter the first results on that condition before firing off the second query, or do the join "backwards" if you think that's going to be better / there's an index available for that. > Their section on reliable replications can be resolved by keeping a log of the changes made to the database rather than a log of the queries. You'd need to define a representation of that log, which ends up being equivalent to writing a NoSQL datastore. > I don't understand how their NoSql model of "Download the latest doc during a change" is any better. It basically becomes "last writer wins" in a distributed system. You at least get consistency. And NoSQL gives you the option of building something better like e.g. Riak does. > The downgrading to NoSql argument makes little sense to me as well. If your frontend is sqlite then you can have Postgres, or MariaDb in your backend. You need to account for having different queries on the front and the back, but the backend is presumably also operating under very different constraints than the frontend. This is kind of the same as the log replication problem - you want the protocol for what's replicated from frontend to backend to be something simple that you could implement on top of approximately any backend datastore. "SQL queries" aren't that, whereas a simple K/V store protocol could be (e.g. MySQL actually implements the memcached query protocol, or did at one point - that would be very hard to do the other way around). > I'd actually make an argument that NoSql databases are a type of relational database heavily optimized towards not needing to do joins Um, WTF? Can you flesh this out at all?
- YarickR2 5y agoHeadline is the question, the answer is "it doesn't"
- sonjaqql 5y agoThe article didn't answer the question for me of how to query for the absence of a property, without getting an entire document collection first. I was really hoping there was a secret to search against null or undefined. Is there a NoSQL solution that does allow for such queries?
- lokedhs 5y agoWith CouchDB you can do it. You create a view filters on the absence or prescense of that field. I'm not sure if that answers your question, as I don't know how other nosql databases handle indexing.
- sonjaqql 5y agoI'll investigate the constraints of CouchDB, thanks!
- sourcesmith 5y agoIf you want to include Postgres jsonb columns in that then a partial index of the expression of a NOT of the jsonb contains jsonb operator.
- mutagen 5y agoThe slightly inflammatory title makes more sense in the context of their database, which sits on top of PouchDB, with adapters for IndexedDB in the browser and a variety of stores on the server side. Based on the series of blog posts / documentation opinion pieces that have been posted so far, I'm quite interested in playing around with this despite being mostly in the relational SQL camp. Everything I've read is thoughtful, well reasoned, and rather practical and the author is exploring a rather interesting problem space. I'd love to see a mashup of RxDB[1] and absurd-sql[2] that brings a distributed SQL datastore to the browser. 1) https://rxdb.info/ https://rxdb.info/ 2) https://github.com/jlongster/absurd-sql https://github.com/jlongster/absurd-sql
- adanto6840 5y agoHearing this repeated, and the comments, makes me start to feel a little bit old. Haven't we re-hashed this out, thousands of times, at this point? There are use-cases for both! There are certainly mis-uses of both, too! DB admins are a real thing; I haven't met many, but the few I have, they are worth their weight. Truly great programmers, I feel the same. They are not mutually exclusive -- solve your domain problem, iterate, and improve from there.... It's that "simple".
- j-pb 5y agoThe author ignores a huge amount of research done in the last 10 years on query evaluation. A series of filter operations and successive sub-selects can't reach the same performance a worst case optimal multiway join can reach (and not even the performance of a series of two way joins that are smarter than linear scans, e.g. a hash join). Similarly, there is no reason why incremental view maintenance should be possible for NoSQL, but not for SQL. Materialize has shown that incremental view maintenance for PostgreSQL is very possible and very fast. This article also seems to very narrowly limit NoSQL to Document Store, which is extremely narrow minded. All in all it does NoSQL a huge disservice.
- dmpk2k 5y agoCould you elaborate on your second paragraph?
- j-pb 5y agoTL;DR Your intermediary result sets can be significantly bigger than your ouput set. By instead of combining whole relations but performing a depth first search through the variable assignment space, while peforming constraint propagation, you can achieve a runtime that is much better than what classical pairwise joining systems can achieve. Since the proposed solution for joins in RxDB is an especially naive kind of pairwise join, it will be especially bad. https://cs.stanford.edu/people/chrismre/papers/paper49.Ngo.pdf https://cs.stanford.edu/people/chrismre/papers/paper49.Ngo.p...
- catwell 5y ago> Yes, there are SQL databases out there that run on the client side or have replication, but not both. Yes, there are. https://zumero.com https://zumero.com for instance.
- ___luigi 5y agoI think the author didn't cover a lot of fundamental research behind NoSQL. If you are reading this comment, and want to dive into SQL/NoSQL DB fundamentals, I highly recommend checking CMU DB (Andy Pavlo) lectures [1] [2]. [1]: https://db.cs.cmu.edu/seminar2020/ https://db.cs.cmu.edu/seminar2020/ [2]: https://www.youtube.com/playlist?list=PLSE8ODhjZXjagqlf1NxuBQwaMkrHXi-iz https://www.youtube.com/playlist?list=PLSE8ODhjZXjagqlf1NxuB...
- DrBazza 5y agoSQL = structured query language. NoSQL typically means non-traditional non-relational database. How many of those actually have truly unstructured data where SQL would genuinely be useless? It doesn't always mean that the vendor of a NoSQL database shouldn't or couldn't implement SQL on top of it.
- wertgklrgh 5y agoDynamoDB is pretty cool tho. You need to know a few limitations and design your data in a certain way but in turn you get like literally unlimited scaling and will never have to worry about overnight success bringing your services down. No wonder AWS has moved most of it's DB usage over to DynamoDB. For me its Postgre -> DynamoDB if working on an MVP and then maturing or if the access patterns are well known to begin with then DynamoDB. One particular thing i love about DynamoDB is the 1 click global tables. I've yet to see anything similarly easy in the SQL world for going global with a database. (Like one AZ in US and one AZ in EU so that every customer gets a low latency) Also, most of the time if Dynamo is used correctly it costs way less than an RDS instance running Postgre.
- tabtab 5y ago> never have to worry about overnight success bringing your services down. But the chances of becoming a viral sensation are very small compared to the probability of longer-term maintenance headaches you'll likely have under NoSql. NoSql is thus similar to meteor insurance. I suspect egos are not making realistic estimates.
- wertgklrgh 5y agoDynamoDB has way less maintenance compared to MySQL/Postgre running on even a managed service
- nine_k 5y agoNoSQL is just a poor term. It lumps together a number of radically different approaches. Imagine a term like NoCar that would lump together airplanes, bicycles, boats, trains, and scooters, just because they are means of transportation which are not a car. Things like Redis, Kafka, Consul, FoundationDB, RRDTool, git, S3, and plain files are all NoSQL databases of sorts. They all are useful in certain areas, all have very different features and guarantees, and each would be a poor substitute for each other or for an SQL database. (Likely even MongoDB can be useful in some areas, even though I have a hard time imagining that.) I wish the whole "NoSQL" moniker would go away, replaced by a few terms that make more sense. That said, the original article is a good one, laying out the upsides and downsides of a document database, and why it makes sense as a local, per-app database.
- mr_overalls 5y ago> radically different approaches Exactly. Relational databases are fantastic general tools, but various use cases can make more specialized data stores the best choice for a particular job. Document stores, key-value stores, column-oriented databases, graph databases can all be more suitable for a given tasks than relational databases.
- tabtab 5y ago> Document stores, key-value stores, column-oriented databases, graph databases can all be more suitable for a given tasks than relational databases. Can related features be add-ons to RDBMS rather than leaving the RDBMS world for good? Any non-trivial system will need at least some of what RDBMS offer. We should try not to throw out the baby with the bathwater. I'd like to explore specific use-cases for NoSql to see what features they need that RDBMS currently lack and if it's impossible to practically add those features to RDBMS. Dynamism of "structure" is about the only thing I can think of right now, and there are possible remedies, per "Dynamic Relational" mentioned nearby.
- mr_overalls 5y agoAt some point when designing a data store, decisions must be made with regards to memory layouts, integrity, transaction guarantees, ease of replication, etc. These choices typically involve reliability vs performance tradeoffs for various use cases. I mean, you _can_ store graph data in a relational database, but its' not typically easy to insert or query, there are serious performance penalties, etc. Purpose-built data structures will always out-perform generic ones.
- menotyou 5y ago"Everything can be downgraded to NoSQL" I am just loooking on our ERP system which has around 580k tables, 10M table columns and around 5.6M programs accessing these tables. For all practical purposes, I really wonder how this should work using unstructured data.
- rwallace 5y agoAre these actual numbers? If so, you've got me curious! How does it work? Do you really have billions of lines of code? How many of those 580k tables are auto generated?
- menotyou 5y agoThis is a standard SAP system. I just checked the numbers before I put them down here. SAP has been around since the late 80ies and is open source and very little of the code has been officially retired, so the amount of code and number of tables are monotonically growing for 4 decades. Lot's of redundant programs as well. So the code base is huge and the database is of course not completely normalized. SAPs program environment comes with a built in editor and a relational database, and that the metadata of every table is stored in the database and the program is not stored on the file system but in the database as well. (Which makes tools like github superflous. When you edit code it gets a lock on the database). So conveniently in SAP there is a database tables which holds all database tables and another table for all fields. Likewise there is a table for all programs and SAP stores all program lines in another database table. So you just need count the number of entries in these tables. I am not aware of any generated tables, but maybe are some. Maybe 5%. I'd estimate that maybe half of the programs are actually generated code. Whereas "program" probably is not accurate, some of the programs are modules or function pools which are invoked by other programs.
- menotyou 5y agoI made some further research today because i wonderd myself. I think the number of generated programs is higher then my estimation. I assess that the number of not generated programs are in the ballpark of 500k and the number of transparent database tables around 300k. The rest of the are views defined on the database but not transparent table. The numbers are smaller than initially given, but still big enough for not being handled with unstructured databases.
- gerbilly 5y agoNoSQL == NoCare about data integrity and consitency concerns. NoSQL allows you to start coding fast (no need to learn any awkward 'legacy' concepts). More importantly it allows you to kick the can down the road for all those problems old people told you about, but which you don't really believe exist. If your project is a one off, NoSQL could be a good choice. But if the data has any value, or especially if it is expected to outlive the original application it was coupled to, SQL is the way to go.
- sethammons 5y agoAn argument I heard the other day: lots of people are realizing the benefits of typescript. Those same people will eventually realize a table schema will make their lives similarly easier.
- toolslive 5y agothe argument is correct, but alas: "eventually" can take a long time. But types and NoSQL can be friends too.
- DonnyV 5y agoBeen using MongoDB for 7yrs now. Best thing I've ever done. NoSQL db like MongoDB cuts down on your joins a lot! Also taking business logic out of the DB and putting it all in the application layer makes things way easier with versioning and sharing code across projects. I love that the mongo database is just used as a black box. It saves data and it queries it, nothing more. This part is rarely talked about. The installation and having multiple versions running side by side is so easy. You can't beat just copying a folder of exe's and then run it and it works. I have systems that have been running for years and have never lost data. I don't use a RDB unless the client demands it.
- ramraj07 5y agoThere’s a reason pretty much any real tech company doesn’t use nosql except in very clear explicit cases. Every single “problem “ you mentioned is not a real problem to begin with or also exists on both sides and you didn’t even realize you were dealing with them. Like migrations.
- randomrubydev 5y ago> There’s a reason pretty much any real tech company doesn’t use nosql except in very clear explicit cases. That is such a bold claim that is obviously not true.
- ramraj07 5y agoCan you give a single example of a major tech company using nosql as its primary db for any application?
- deleted 5y ago[deleted]
- jd_mongodb 5y agoMongoDB hss 26500 customers worldwide (and we are one NoSQL vendor). These customers include: Bosch : https://www.mongodb.com/customers/bosch https://www.mongodb.com/customers/bosch HSBC : https://diginomica.com/hsbc-moves-65-relational-databases-one-global-mongodb-database https://diginomica.com/hsbc-moves-65-relational-databases-on... SEGA Hardlight : https://www.mongodb.com/blog/post/sega-hardlight-migrates-to-mongodb-atlas-simplify-ops-improve-experience-mobile-gamers https://www.mongodb.com/blog/post/sega-hardlight-migrates-to... HMRC : https://www.mongodb.com/blog/post/mongodb-microservices-help-hm-revenue-and-customs-cut-wait-times-in-half-and-save-gbp8-2m-in-operational-costs https://www.mongodb.com/blog/post/mongodb-microservices-help... DWP : https://www.mongodb.com/customers/department-for-work-and-pensions https://www.mongodb.com/customers/department-for-work-and-pe... Liberty Mutual : https://www.mongodb.com/blog/post/liberty-mutual-iac-mongodb-atlas https://www.mongodb.com/blog/post/liberty-mutual-iac-mongodb... MetLife : https://gigaom.com/2013/05/07/with-300m-earmarked-for-tech-innovation-metlife-wants-to-remake-insurance/ https://gigaom.com/2013/05/07/with-300m-earmarked-for-tech-i... There is a more complete list here : https://www.mongodb.com/who-uses-mongodb https://www.mongodb.com/who-uses-mongodb That list is limited is just the customers that are willing to be public references. Every mature NoSQL vendor has a similar list.
- clpm4j 5y agoUse the right tool for the job. SQL databases are not always the right tool. NoSQL databases are not always the right tool. If you think an RDBMS is the best choice for every problem, then maybe you should stop and ask yourself if it's just the only tool you're comfortable with. There is a reason that the number of NoSQL options continues to expand - if nobody was using them successfully then they probably wouldn't continue to grow.
- tabtab 5y agoA key aspect that RDBMS lack compared to NoSql is dynamism. Dynamic Relational is a draft RDBMS idea that intends to fill that niche. Unlike the NoSql movement, it tries to keep as many "traditional" RDBMS idioms as possible when adding dynamicness. It also allows one to "lock down" the schema incrementally as projects mature. https://stackoverflow.com/questions/66385/dynamic-database-schema#46202802 https://stackoverflow.com/questions/66385/dynamic-database-s...