3 ms·
We use it for our database engine, as the nature of our data does not work in a SQL table. I work with DNA sequencing, and we store our data in a sample metadat
by jakereps 9y ago
We use it for our database engine, as the nature of our data does not work in a SQL table. I work with DNA sequencing, and we store our data in a sample metadata, feature metadata, and sample-by-feature count table.
So our sample metadata has the generic metadata about each sample (where it's from, when was it collected, etc.). Our feature metadata is the particular feature metadata (DNA amplicon sequence variant, taxonomy, all seven taxa levels pre-split). These would work perfectly fine in a standard SQL database, but the problem is when we get into our sample-by-feature table/collection. Currently, we have over 400k unique features and tens of thousands of samples. If we were to try and map the frequency of each feature in our samples in a standard row/column relational database table, our table would be un-usably large (or just entirely not work as PostgreSQL, MySQL have a hard column limit). Thus, the benefit of MongoDB's document system fixes all the issues this causes.
Each of our documents in the sample by feature table can consist of ONLY the counts of the features that were in the sample, ignoring all the other N-thousand features that may be present in others. Building a pandas DataFrame from a query of this data will just fill the absent features with NaN, and we can just fill those NaN's with zeroes and carry on working with our target subset of the database.
- smt88 9y agoYour entire response seems to be about why you chose document storage instead of relational data. Relational data makes sense most of the time, but not always. The question, again, is why you'd use MongoDB even for document storage when there are so many alternatives that are safer and better-engineered. Even Postgres has a JSON data type that, by some reports, works better than Mongo and is much safer to use. Another thread here[1] has gone into some detail about alternatives. 1. https://news.ycombinator.com/item?id=15308214 https://news.ycombinator.com/item?id=15308214
- jakereps 9y agoIf that's the case, then it's probably purely marketing. I had no idea PostgreSQL offered a comparable option.
- smt88 9y agoSomething I don't think is mentioned elsewhere: other databases have shipped with Mongo compatibility. This means you can migrate from Mongo to something better without changing your code. The most mainstream and interesting of these is Cosmos: https://docs.microsoft.com/en-us/azure/cosmos-db/connect-mongodb-account https://docs.microsoft.com/en-us/azure/cosmos-db/connect-mon... (Cosmos is actually one of the most interesting DB products available for production right now, in my opinion.)
- deepsun 9y agoYep, I benchmarked it myself, JSONB column type in PostgreSQL worked faster for me than same data in MongoDB (both reads and writes).
- btown 9y agoJSON wasn't stable in Postgres until 2012, whereas the "web scale" memes for Mongo started after its initial release in 2009. Mongo indeed had better marketing and a first mover advantage, and it never let go.
- chx 9y agoAnd it wasn't usable until practically 2015 (2014-12-18 if you so want).
- cervo 9y agoEarly versions of mongodb locked the entire database server for every single write. Later versions locked the entire database for every single write. Only since Mongo 3.0 does MMAP only lock a single collection for a write and is WiredTiger available to offer you MVCC. And even in the MongoDB world materials it advertised now you can use more of your hardware.... Granted you could always shard, though that gets very expensive very quickly. Meanwhile postgresql had MVCC all along and its write speed was always faster than mongo on a single server basis. And you always could serialize data that changed into text fields as xml/json/csv/some other format. Mongo was mostly great marketing and a lousy product. Over time it has gotten much better. Mongo 3.2 is a way way way better product than mongo 1.x. But marketing did capitalize on a lot of hype which made no sense. Developers enjoyed just serializing their objects (the ones who didn't think of doing this into a text column in a relational db) and for the ones who did not need 'web scale' they had no idea about the severe concurrency issues introduced by mongo's writes. And for the rest there are plenty of posts about hacking around it or switching from Mongo to something else.
- tzmudzin 9y agoMaybe I misunderstood your problem, but how about a relational DB model with: - FEATURES table (~400k rows), primary key: Feature_ID - SAMPLES table (tens of thousands), primary key: Sample_ID - OCCURRENCE table (what you observed), with the fields Feature_ID, Sample_ID, Occurrence_Count. This is pretty much a standard solution well proven in data warehousing...
- gaius 9y agoIf you process this kind of high-throughput data in R then yep, you get it as 3 data.frames as you describe, it's nicely tabular and easy to "join" by just matching index numbers.