3 ms·
We have mostly moved off mongo at this point, there remains a single tiny mongo cluster running with a handful of collections that aren't worth the time investm
by mbell 6y ago
We have mostly moved off mongo at this point, there remains a single tiny mongo cluster running with a handful of collections that aren't worth the time investment to move at the moment. Almost everything moved to PG. The abstract issue we had with mongo is the purported 'best practices' with using it were seemingly in conflict with its actual implementation. I should note that these are issues across the last ~5 years so it's likely some of this has changed, it's also likely my recollection of the details are not perfect.
Mongo pushes the idea of keeping related data in a single document. So if you have a hierarchy of data, keep in all in a nested document under a 'parent' concept, say an 'Account'. The problem with this is that there is a document limit of 16MB and key overhead is high. At one point we had to write a sharding strategy where by data would be sharded across multiple documents due to this limit. This also broke atomic updates so we had to code around that. We also ran into a problem where for some update patterns, mongo would read the entire document, modify it, then flush the entire document back to disk. For large documents, this became extremely slow. At one point this required an emergency migration of a database to TokuMX which has a fast update optimization that avoids this read-modify-write pattern in many cases, as I recall it was something like 25x faster in our particular situation. This same issue caused massive operations issues any time mongo failed over as the ram state of the secondary isn't kept up to date with the master so updates to large docs would result in massive latency spikes until the secondary could get it's ram state in order. In general we just found that mongo's recommended pattern of usage just didn't scale well at all, which is in contrast to its marketing pitch.
I think at one point we had something around 6TB of data spread across 3 mongo clusters. After migrating most of that to PG or other stores and reworking various processes that could now use SQL transactions and views, the data size is a small fraction of what it was in mongo and everything is substantially faster. In one extreme example there was a process that synced data to an external store as the result certain updates. Because we couldn't use single documents and had no cross document update guarantees we would have to walk almost the entire dataset for this update to guarantee consistency. It got to the point that this process took over 24 hours and we would schedule these update to run over a weekend as a result. With the data moved PG, that same process is now implemented as a materialized view that takes ~20 seconds to build and we sync every 15 minutes just to be sure. Granted this improvement isn't just a database change but rather an architectural change, however mongo's lack of multi-doc transactions and document size limit are what drove the architectural design in the first place.
Then there are bugs, of which there were many, but the worst of which was a situation where updates claimed to succeed, but actually just dropped the data on the floor. I found a matching issue in mongo's bug tracker that had been open for years at that point. Ultimately I just can't trust a datastore that has open data loss bugs for years, regardless of its current state.