6 ms·
NoSQL: The Baby and the Bathwater
- personomas 4y agoStructure is important, that's where NoSQL fails from the get go. We should stop investing so much time into NoSQL and look more into combining SQL and Graphs
- emodendroket 4y agoThat seems a little like saying there's no reason for anyone to use a language without GC. Sure, for most applications it is fine.
- trashtester 4y agoThat's a good analogy. SQL RDBMS's, just like Java and other languages with GC have huge utility in a multitude of enterprise software situations. These probably make up the majority of software systems and lines of code out there. Then there are a few domains that NEED other solutions. In terms of CPU cycles these may be consuming more resources. These are operating systems, game engines, ML libraries, search engines, social media platforms and huge webshops like Amazon. But to store all data in Spark or MongoDB just because that's what <insert Cargo Cult idol> is doing makes about as much sense at programming everything in C/C++.
- winrid 4y agoYou can use NoSQL with a schema and have strict enforcement. What the NoSQL db does is make storing objects easy with less boilerplate or abstraction.
- 9dev 4y agoHow is NoSQL with schema validation and strict enforcement anything but SQL with extra steps?
- dgb23 4y agoThe schema restrictions you can express in SQL are both insufficient and not flexible enough even for relatively trivial things, except you are either fine with storing and retrieving things in a completely different shape than you actually need, or if you are fine with constantly migrating your schema. That's not a sufficient argument against adopting SQL for something, but other solutions are certainly _not_ just SQL with extra steps. Not everything maps nicely onto flat tables with predefined slots.
- hugofirth 4y agoYou may be interested in https://www.gqlstandards.org/ https://www.gqlstandards.org/ and https://www.iso.org/standard/79473.html https://www.iso.org/standard/79473.html TL;DR the ISO standards committees behind SQL are working on bring graph query languages and SQL databases closer together. Obviously there is work to be done at the storage/query planning layer after that, but I’m hopeful once the surface exists more widely that will drive more work in those areas.
- threeseed 4y agoYou would think after all these years people would stop using the term NoSQL. They are just databases. Many of the so-called NoSQL ones like MongoDB or Cassandra can support schemas, transactions, strong consistency, joins, secondary indexes etc. And SQL databases like PostgreSQL support schema-less data structures. And with Presto, Spark SQL etc you can use SQL with almost any data store.
- reidjs 4y agoJust because they support those things does it mean it’s straight forward to use them Our team had a lot of trouble trying to map highly relational data to a noSQL database (mongo) It could have been a failing of our team, but I also think it just made our lives way harder. We’re on Postgres now and a lot of issues have faded away.
- revscat 4y agoSame experience here. We had a small-ish application that was originally built in top of MongoDB. Once it made it into production and started to see some success, it became quickly apparent that the schemaless design caused problems that an RDBMS would have solved. It was decided fairly quickly to remove the MongoDB underpinnings and migrate everything to Postgres. It was the right call, and the final nail in any remaining affection — and interest — I had for NoSQL stores.
- collaborative 4y agoSure SQL DBs can have all the advantages of NoSQL DBs and then some... but they will never have a lower price tag. And that's because NoSQL DBs use very, very little CPU. All that's needed is giving up reliance on SQL. Design your NoSQL "schema" with this in mind, and you're golden
- winrid 4y agoVery little CPU depending on sorting, indexes, etc... Have a lot of indexes on a collection? Mongo will eat 5-10ms of CPU time per query even with cached query plan stats just to start executing the query. So no different than PG here. What you're getting at is Joins, but I haven't seen a company that didn't end up doing joins with Mongo at some point. Or, they do it in the app layer, in Java/Node/PHP which requires more CPU than it would in the DB. Also, wanna store an array in a document in Mongo? Every time you add to that array, the whole array is replicated to your secondaries. That eats a lot of CPU.
- collaborative 4y agoBut you are describing incorrect ways of using NoSQL. NoSQL requires getting the schema right from the get go. Many people don't like it because they are used to relying on SQL (or a language like you mentioned) to smooth things out. NoSQL requires you to ask yourself hard questions - what specific queries are going to consume this data and then model the data as opposed to in SQL where you just store the data and worry about querying it later. This produces efficient queries that ultimately don't need as much CPU
- trashtester 4y agoSome claim the opposite, though. Even in this discussion. That NoSQL is needed when you need to change the schema in a production system. Anyway, most systems that last more than a couple of years tend to change over time, meaning the schema is likely to need extensions.
- 4y ago
- roenxi 4y agoI don't like NoSQL databases. "NoSQL" should be manna from heaven. It may be impossible to create a less pleasant language than SQL. It isn't composable, it isn't internally consistent, it isn't easy to parse, it claims to be declarative but the ordering of the clauses is completely rigid, it fights every attempt at writing testable or maintainable code. It is hard to read. It is hard to programatically generate. Everyone seems to be trying to develop systems that mean they don't have to write SQL code. Add all that up and NoSQL should be fine. But then the actual NoSQL databases let people throw out all the concepts that make SQL so sticky. Schemaless databases are in the same class as goto statements - there are people I trust to use them to do amazing things. The projects I inherit to maintain are never written by those people. Transactions and consistency guarantees are necessary to get a group of programmers together writing a reliable application. The relational model is the best idea to come out of database research. I wish NoSQL was a postgres extension. Instead we get MongoDB advocates. Bless them, but PostgreNoSQL would be so much better for all the use cases I have than MongoBD.
- emodendroket 4y agoI think the right way to think about NoSQL databases is you have an application where traffic is expected to be heavy enough that it's worth throwing out what SQL gives you for free and having a highly optimized solution where you deal with those problems yourself. Of course, it was a big enough trend that many people jumped on it without having a practical use for it. Arguably psql does offer a NoSQL-like experience with JSON columns.
- roenxi 4y ago> Arguably psql does offer a NoSQL-like experience with JSON columns. I hope not. JSON columns purge all the good things a relational database offers and keeps the SQL. It is the worst of every option. That really should be the opposite of the NoSQL experience. Disbarred by the fact that SQL is involved.
- zztop44 4y agoJSON columns are very useful in some circumstances but you shouldn’t use them as a replacement for a database schema, and if I’m honest I’m yet to see anyone truly suggest that.
- benrutter 4y agoI'm surprised to see so many opinions on what is "objectively" the best database paradigm. My incredibly boring take: databases are such a generic tool that it really depends on your use case, there are times when noSQL is stand out the best choice, and lots more when it isn't- the real pain point comes from people blindly cargo culting into a db paradigm without considering their use case.
- tgv 4y agoYup. I've got a (legacy) postgresql database with time series: one timestamp plus measurement per row. And there's a useless third column too. It's awful, and sluggish. For a newer project, where we basically use one object, I've chosen a NoSQL database. Get the object, edit it in the browser, put it back. Done. No need to update relations, ORMs or any of that. But: that won't fly for more complex projects. So I agree: pick the right tool for the job.
- nhourcard 4y agoHi tgv, curious to hear why you picked NoSQL for your time-series use case?
- tgv 4y agoIt's not a choice that's guaranteed to suit your use case, but storing a time series as a large set of rows is not performant. Perhaps I should have mentioned that our time series data is "write once": we record and store it, but it doesn't get altered. And the time series are all sent to the browser, which does the displaying, filtering and analyzing, so there's no point in processing (our particular) time series in SQL, if that were even feasible. So, basically, because it is good enough, and better than the available alternatives.
- bullen 4y agoExplicit schema is the most important feature to drop. You cannot have a live service that is interrupted by database modifications. My solution is to use JSON. But the most important feature to add is HTTP as transport and async-to-async (both client and server needs to only use a thread when they are doing work, zero idling): To scale a distributed database across continents: 2000 line distributed DB: http://root.rupy.se http://root.rupy.se
- trashtester 4y agoIt may be the most important feature to drop, for some. It's certainly the most important feature to KEEP for others. Schemas allow you to keep the more critical business relationships consistent. For variable data elements, you can always use JSON columns for that.
- regularfry 4y ago> You cannot have a live service that is interrupted by database modifications. So don't do that, then. Designing database migrations to be non-breaking is part of the game and if you're not doing it, you can't claim to understand the technology you're replacing. Not having an explicit schema doesn't mean you don't have to think about your schema. It just means you've chucked out all the tooling for keeping it sane.
- throwawaaarrgh 4y agoMost databases are still kinda crap. Either it's hard to scale them or maintain them, hard to know the right way to use them, or they strip out useful functionality for the sake of simplicity, forcing reinventing the wheel. We haven't got much actually novel design since like 2006. There's some people trying to bolt on top of/around/under existing databases for backwards compatibility or to avoid reinventing, but not really much actual novel design for a whole database. Computer Science research seems to either be fully academic, where research into practical solutions doesn't happen, or it's limited to private companies trying to solve one business problem. Very little actual advancement of the state of the art, or solving of long standing limitations.