5 ms·
> By far the most consistent mistake was choosing a non-relational database, when your data was strongly relational. Mongoose's ODM made this mistake surprising
by MikeKusold 9y ago
> By far the most consistent mistake was choosing a non-relational database, when your data was strongly relational. Mongoose's ODM made this mistake surprisingly easy to make, which led to issues down the line.
This mirrors my experiences with Mongo as well. The vast majority of data is relational. Mongoose allows people to make the mistake of structuring their data relationally. But all you are doing is pushing all the joins to the web server (which in Node land is single threaded and compounds the mistake of choosing NoSQL).
There may be some microservices that Mongo is great fit for, but it should not be your core data store.
- mrtksn 9y agoIf data needs to be processed, an option would be to use MongoDB for collecting the data in bulk and later decide how you need to structure it for your needs. When it comes to "read" the data, you read it only from processed database that can be anything. Since you can do your processing completely independently of your web server, your are not necessarily pushing your computational load from the DB to the web server.
- emodendroket 9y agoIt that point it's harder for me to understand why I'm not just writing out JSON files or something.
- mrtksn 9y agoYou definitely can but MongoDB provides convenience over storing and managing JSON files on you filesystem by your own efforts.
- emodendroket 9y agoSure, but it is also another service to keep running. Depends on what scale you're operating at.
- mrtksn 9y agoAny RDBMS is just a toolset to manage files on some filesystem. If you are dealing with some config files you can spare yourself the trouble and just load them from the filesystem. NoSQL or classic RDBMS, these tools give you ways to manage your data on the filesystem(or in the memory?). You can spend time to write a code that processes files in a folder or you can use the software that does it for you. You can write a code that relates two files in a folder or you can use a software that already can do it and you can just ask that software for the outcome.
- emodendroket 9y agoI don't really have to write that code though; what language doesn't already have file-manipulation libraries?
- mrtksn 9y agoWell I guess some people wonder why people are using ready software when you can achieve the exact same thing with a text editor and a compiler.
- emodendroket 9y agoThat's not the same thing at all; you're suggesting an approach that probably takes more code and definitely takes more setup.
- mrtksn 9y agoYou are underestimating the complexities of file management
- emodendroket 9y agoAgain, depends on scope and scale. If you have a handful of items to process it's not a big deal.
- deleted 9y ago[deleted]
- lowbloodsugar 9y agoNo, I really can't think of any situation where installing, managing, updating, maintaining, and crying over a mongodb cluster would be easier for a flat-file datastore than say, a file system, or S3. In fact if one is on AWS, then nothing beats just dumping them in S3 and processing them as necessary. If I'm a startup and I really wanted some querying then PostgreSQL looks great, with Amazon handling the maint and backups.
- bdcravens 9y agoYou can even query flatfiles in s3 using SQL via AWS Athena.
- flavio81 9y ago> If data needs to be processed, an option would be to use MongoDB for collecting the data in bulk and later decide how you need to structure it for your needs. OR, i can open a file stream, serialize my data to JSON, entity by entity, and dump all to a file. The good old file.
- mrtksn 9y agoreplied to the same question here: https://news.ycombinator.com/item?id=14805812 https://news.ycombinator.com/item?id=14805812
- threeseed 9y agoBecause we want to do partial updates, searches, indexing etc. Your position can be applied to all databases. Why not abandon them all and just use CSV ?
- bdcravens 9y agoYou could always do that, dump the data to s3, and use Athena :-) https://aws.amazon.com/athena/ https://aws.amazon.com/athena/
- pritambaral 9y agoAnd then be locked-in to Amazon. From ubiquitous files to ... a single service provider.
- aidos 9y agoThen again, if you're going to push into Postgres you might as well store the temp unstructured records there too
- mrtksn 9y agoWell, that's also an option :) But I believe this is a relatively new option. Personally, I never got into the hype with the NoSQL and for me the only use is as a convenient API for storing JSON files during prototyping so that I can later decide how my data should be structured. Definitely tech debt for the future but fast iterated design that needs the database stuff fixed has more chance for success than the rigid design with the perfect data architecture.
- scarface74 9y agoBut all you are doing is pushing all the joins to the web server (which in Node land is single threaded and compounds the mistake of choosing NoSQL). Why is pushing the joins "to the web" inherently bad? It's a lot easier to scale web servers than database servers.
- bdcravens 9y agoJoins are typically a tiny intersection of 2 data sets. Multiply that by the number of web requests. Moreover, a query planner is far more optimized than the code you'd write on your web server.
- kazen44 9y agoalso, scaling databases is not terribly difficult, it is just very "clumsy". Still, having a proper transaction manager and query planner will give you far better performance compared to doing it on webservers. Also, aren't bsically moving the complexity of the database to the webserver? Which makes your webserver far more complex and harder to scale in comparison. Especially if you need ACID. (which is a must if you are doing anything with data storage imo).
- scarface74 9y agoYou don't always need ACID. Sometimes eventual consistency is good enough. How would it make the webserver itself more complex? The application itself should just be able to run on n servers.
- nulagrithom 9y agoSimplified, a JOIN on an SQL server will take advantage of its indexes to grab the data from the disk that matches the specified criteria. SQL databases are built for this sort of thing, and they're really freaking good at it. Sending that same JOIN to the web server means, at the least, shipping everything from both tables matching the criteria and matching up those two potentially huge blobs of data with each other on the web server. Doing this in an even partially efficient manner will split your query logic between DB queries and other code. Also, it's not unlikely that you'll be doing this JOIN operation in a language that maybe isn't great at optimizing that sort of CPU intense behavior on its own, so you're probably going to need to drill in and optimize on your own (vs letting SQL do its magic). Plus, we're going to need to store all this data in memory while we match it. So now our web servers, while more easily scaled than a DB, are actually pretty damn big boxes with decent memory and CPU, and that's expensive to scale.