3 ms·
I do like the "just throw whatever you want in this document" approach that Mongo has.
by colonelpopcorn 8y ago
I do like the "just throw whatever you want in this document" approach that Mongo has.
- StavrosK 8y agoI've regretted that every time I've done it, with the usual "why is this document not like the others" failure case. Nowadays I use schemaless databases/fields extremely sparingly. Relevant talk: https://www.youtube.com/watch?v=BN8Ne2JCGBs https://www.youtube.com/watch?v=BN8Ne2JCGBs
- tyfon 8y agoWe are using azure documentdb for a project at work and it already has like 3-4 structural versions of the documents in the db that has to be accounted for by the applications and reporting tools, the latter of which I am working on. It is somewhat similar to Mongo. It's starting to become really painful and I'd be extremely happy if they moved it to a relational database.
- StavrosK 8y agoThat's been my experience too. Very rarely is your data model non-relational and schemaless, and if it isn't, you end up having to manage all the relations and cascades and everything yourself, which ends up being a huge pain.
- e12e 8y ago> Very rarely is your data model non-relational and schemaless This of course completely false - which is why we use filesytems and files, more than we use databases ;-) But what I think is less clear and obvious - is when is a document database better than files and the file system? This is kind of the question, are maildirs/mbox any good at storing mail - and are they any good at searching mail? And is sql better at either? This isn't clear cut - there are things like dbmail[d] that just uses postgres and layer imap on top/in front - while a lot of servers prefer some form of file storage augmented with custom indexing. [d] http://dbmail.org/ http://dbmail.org/
- StavrosK 8y ago> This of course completely false - which is why we use filesytems and files, more than we use databases Filesystems and files are good at a lot of things, but not really great at anything. They're a general-purpose abstraction, but the relational model is much more specific.
- e12e 8y agoOf course. I agree with the points made here and in sibling threads about schemas and sql databases. But I also think it's hard to dispute the fact that most data lives in files and file systems. And I think developers too often compare various no(t only)sql solutions with rdms' - without really considering "file-based sql" (sqlite) - or simply file system/DAV based storage (webdav, s3 and work-a-likes, local files, nfs etc). And again - I think it can be easy to see some examples where files are great (store the images for an image gallery), or where relational model is great (transaction data). But I think it can be much harder to see where file storage, or file storage augmented with some form of index - is beaten by a nosql dbms.
- msbarnett 8y ago> This of course completely false - which is why we use filesytems and files, more than we use databases ;-) But consider: there is no application in existence that attempts to work with the contents of every kind of file that can possibly exist on a filesystem. The only kinds of information that any application can reasonably aggregate across all files are those which conform to the standard metadata schema the filesystem applies to all files. Creating an application that aggregated anything else across all files is essentially impossible -- no sane human would attempt to write an application whose functionality depended on them writing parsers for every single file format that could possibly be found in a filesystem. And yet, for most uses of MongoDB it is acting as the datastore for a single application, which in order to do any number of things normally demanded of applications (queries, aggregations, CRUD, etc) must be able to read every variety of document that anyone has ever thrown into the datastore, whether those variations were planned (lets add the following fields to all new Comment documents) or accidental (whoops, a bug led to 3 days worth of Comments missing one or more of the following fields). In the latter scenario, the "it's like a filesystem, store whatever you want!" actively works against the controlling the complexity of the application, to the programmer's great detriment.
- jerf 8y agoEven when I'm storing data in a schemaless fashion, I always use some sort of strongly-typed mechanism for generating the JSON that provides guarantees about the nature of the JSON I am storing. At the very worst I end up with multiple well-defined versions of the "schemaless" content. "Schemaless" here just means the default relational mechanisms don't work well for my use case for some reason. "True" schemalessness is a lie. You can't even hand arbitrary blobs of JSON to human beings and get consistent interpretations of them. There is always a schema. I'm not dogmatic about where the schema lives; relational DBs are great when your data fits them but there are times you need an "exception" to your relationships for some chunk of data. I'm fine with the schema being "this particular strong type in a programming language as rendered by this JSON renderer"; that's a fine schema too. But some schema always exists, and no matter how you slice it, you can not escape from issues of dealing with schema changes. Thinking you can just ignore the problem and just throw whatevs into the database is simply one of your options, one which may gain a lot of very, very short-term convenience for a loooot of long-term pain. As with all cost/benefit questions, there are times where that is the right choice... but it's a lot less often than the old NoSQL hype presupposed. (I've used this option when extracting things from an existing, stable DB and doing an ad-hoc one-off reporting analysis using a temporary database. I don't have to worry about breaking the stable DB when I'm only reading from it, and I can fluidly experiment and develop the one-off report easily. But it's really a rare use case.)
- StavrosK 8y ago> "Schemaless" here just means the default relational mechanisms don't work well for my use case for some reason. Agreed, but you're basically doing things as they should be done. Not using the schema in a relational database because of valid reasons, still validating the schema, etc. I'm not sure if you're agreeing or if you meant it as a different viewpoint, but it sounds like we agree. > But some schema always exists, and no matter how you slice it, you can not escape from issues of dealing with schema changes. This is why I hate it when people say "oh I use MongoDB because I don't want to have to bother with migrations". No, your schema is always going to change, you're just going to migrate things in an ad-hoc manner instead of with a well-thought-out, established process.
- 33degrees 8y agoIndeed: I recently ran into a bug where a field was inserted as a string in some cases and an integer in others. A simple fix, yet more work than I would have had to do had I just configured the field properly in the first place.
- colonelpopcorn 8y agoIt beats the heck out of serialized JSON files. However, for any application that needs data with relationships (most all of them), it's a no brainer to use SQL for your query language and DB engine.
- tibbon 8y agoIt’s always felt then like “easy in, hard to get out”