4 ms·
Hi javiermaestro! All benchmarks have been done on AWS i2.xlarge instances. They have 4 vCPUs, 30GBs of RAM and 800GB local SSDs. It is true that we are not d
by gortiz 10y ago
Hi javiermaestro!
All benchmarks have been done on AWS i2.xlarge instances. They have 4 vCPUs, 30GBs of RAM and 800GB local SSDs.
It is true that we are not doing an apples to apples comparison on the 100GBs set. I think is just the opposite! We are helping MongoDB by giving it three machines (and therefore 90GB RAM!). This is specially true when the benchmark is using indexes because in this case the index can be always on RAM.
MongoDB is very fast when it has to retrieve a single document because, by design, it has an amazing spatial locality (the whole document is usually on the same page). But this feature is a weakness on aggregation queries, as they usually only care about a small subset of the document. ToroDB Stampede change that by storing your data on a relational way. Of course, as you said, MongoDB performance is horrible when it has to fetch documents from disk, but even if the documents are in memory, the same effect is expected (on aggregation queries) when data has to be move to the CPUs caches.
- javiermaestro 10y agoLOL I re-read the article and found the specs. I really read it but somehow managed to skip the paragraph or something :-? Anyway, my point stands. I'd use a single instance in which the dataset fits in memory, just for completeness. Then, you can compare one mongo with the full dataset to one stampede. As you said, the aggregated data will still make Mongo suffer, but it will be a better comparison. I still like the 3-shard setup, though. It's also a good reference point.
- ahachete 10y agoSo far we have benchmarked situations where dataset > RAM or >> RAM. I think it is an interesting point to also analyze the case when dataset < RAM, to see how efficiently both systems manage the caches, query planning etc. Stay tuned and thanks for the suggestion! :)