3 ms·
Actually, having been through their training, and dealt with their consultants. Their company culture is the problem. They had a pure-developer centric mindset,
by edgan 10y ago
Actually, having been through their training, and dealt with their consultants. Their company culture is the problem. They had a pure-developer centric mindset, and no operations mindset. So at the time they had no good way to do backups. Their more modern solution is backups that they manage for you, which is even more crazy.
- Fiahil 10y agoYes. Mongo is an operational nightmare. I mean, who in their right mind would use a product where the recommended solution[0] to resync a stale replica (which happen all the time after a long netsplit) is to "perform an initial sync". Which, of course, mean "please remove everything and type this command". Crazy. [0]: https://docs.mongodb.com/manual/tutorial/resync-replica-set-member/ https://docs.mongodb.com/manual/tutorial/resync-replica-set-...
- bogomipz 10y agoIndeed, and the only practical means of performing a compaction is to "rm -rf" the data directory and let it resync from another replica set member. This is not documented of course.
- Fiahil 10y agoThis is documented (okay, a little "hidden")! It's written black on white in their documentation (see the link above): > A replica set member becomes “stale” when its replication process falls so far behind that the primary overwrites oplog entries the member has not yet replicated. The member cannot catch up and becomes “stale.” When this occurs, you must completely resynchronize the member by removing its data and performing an initial sync. > MongoDB provides two options for performing an initial sync: > Restart the mongod with an empty data directory and let MongoDB’s normal initial syncing feature restore the data > Restart the machine with a copy of a recent data directory from another member in the replica set. Note: the second option is not a real option when you're dealing with a 700GB database. By the time you finish the copy the oplog will be too big anyway. Thus, making all these steps completely pointless. And that's why it's so bad. They even acknowledge the "correct" solution is to rm your data and resync.
- bogomipz 10y agoIt's documented for the use case of "stale replica" but I was referring to the use case for when you want a compaction to reclaim disk space. For that they recommend the db.repairDatabase() option but that requires you to have twice the size of your db available available on whatever partition your database is on. That was I said "practical." But yes the procedure is the same.