7 ms·
> I would likely not use it as a main source of truth for any application but for a lot of things, it's a good database. I really don't understand this. In wha
by devishard 10y ago
> I would likely not use it as a main source of truth for any application but for a lot of things, it's a good database.
I really don't understand this. In what case is it every acceptable for a data store to lose data? And that's not even MongoDB's only problem: it memory leaks!
- avitzurel 10y agoI have to say that I did not experience a single data loss that was a result of the database misbehaving. Memory leaks weren't a huge issue for us as well. After stabilizing the setup I have to say it was basically a fire and forget part of the stack for us. The role of MongoDB was to act as a fast-insert and aggregation framework for other parts of the system. So, we would insert BIG amounts of data at a time and aggregate it into a K/V store where we pulled the data from. After a while, the setup became to expensive to run, at this point we turned it off for a cheaper solution, but in terms of functionality, it functioned pretty well.
- mack73 10y agoFire-and-forget is awesome for when you do not care at all about your data. I'm sure those types of writes are super (duper) fast. What exactly is the use case for that? Serious question.
- vidarh 10y agoDon't know about the guy you replied to, but e.g. consider any application that regularly crawl feeds, api's etc. where the data is rapidly changing and only a portion of the data is necessary to give good output. There are lots of applications like that where you just need "enough" data to give good results and/or where any loss will auto-heal next time you crawl the original source.
- mack73 10y agoThat makes sense to me. If your MongoDB cluster under preasure will only actually persist 90% of your writes and this is something you anticipate, then MongoDB seems like a good choice, if writing in this style is faster than other nosql systems (that make grander promises about persistance) that is.
- vidarh 10y agoExactly - the important thing is you need to actually understand the risk, and make an informed decision what level of loss is ok to you (and you should understand whether or not it's actually saving you anything - as you say, it makes sense if it is faster; there's no point losing data if you don't gain something from accepting the risk). This is also perhaps the biggest problem with MongoDB: It's fast but unsafe "out of the box", and not everyone will know that when they use it. I think that's a large part of the problem a lot of people have with it.
- devishard 10y agoI understand that some data loss is acceptable sometimes, but you can get the super-fast writes with a store that doesn't drop your data at random (such as Cassandra) so why would you choose the store that does?
- nollbit 10y ago> I have to say that I did not experience a single data loss that was a result of the database misbehaving. How do you know? Was your data checksummed? Did you read back changes after writes to verify what was written?
- devishard 10y ago> I have to say that I did not experience a single data loss that was a result of the database misbehaving. Okay, but numerous people have experienced data loss using MongoDB. You can spend 10 minutes on Twitter and find someone talking about it. > Memory leaks weren't a huge issue for us as well. After stabilizing the setup I have to say it was basically a fire and forget part of the stack for us. "After stabilizing the setup"? So basically you worked around MongoDB's stability problems instead of choosing a solution that didn't have stability problems? > The role of MongoDB was to act as a fast-insert and aggregation framework for other parts of the system. Sure, but there are other solutions that can do this without MongoDB's issues (i.e. Cassandra). > After a while, the setup became to expensive to run, at this point we turned it off for a cheaper solution, but in terms of functionality, it functioned pretty well. Sure, you can build something stable on sand, but it's going to be a lot harder than just building on a solid foundation.
- avitzurel 10y ago> Okay, but numerous people have experienced data loss using MongoDB. You can spend 10 minutes on Twitter and find someone talking about it. Never said people didn't. > "After stabilizing the setup"? So basically you worked around MongoDB's stability problems instead of choosing a solution that didn't have stability problems? No, I haven't worked around the limitations, after tweaking heap sizes, machine sizes and disk speeds the setup was very stable. > Sure, but there are other solutions that can do this without MongoDB's issues (i.e. Cassandra). Again, I DO NOT disagree, there are other solutions and today we are running a completely different setup to replace the same component. > Sure, you can build something stable on sand, but it's going to be a lot harder than just building on a solid foundation. I'm trying to ignore the snarky/attacking nature of your comment. Not sure if this is addressable.
- vidarh 10y agoHere's one: I have an aggregation engine where 99.99something% of the data is cycled out within 3 days, and where the system can be functional again within about ~5 minutes of ingesting new data after a total data loss. There are a lot of applications where your database is not the or even a source of truth, but effectively a big cache that you can either fully rebuild or where rebuilding isn't necessary (the data is too "fast moving" for there to be much point). Here is another: A search engine for classifieds where the source of truth of 99%+ of listings were external feeds that'd get re-crawled at least once a day. If we lost a few million updates, who cared? If it was few enough, we'd just let the normal updates take care of it. If we had a major problem, or needed to do a major update that'd have compatibility implications, we could just do a re-crawl of all the feeds. Here's another one: You're scaling reads by replication, and so the vast majority of your data exists in multiple data stores. You may choose one you consider reliable for the source of truth, and the rest are basically caches, but you may want/need something more capable than a straight up key/value store for various reasons. I believe the vast majority of data stores I've worked with have been ok to suffer a total loss of because most of them are not the source of truth, and the source of truth can be re-queries fast enough for it to not be a big deal to deal with losses in secondary copies.
- devishard 10y ago> There are a lot of applications where your database is not the or even a source of truth, but effectively a big cache that you can either fully rebuild or where rebuilding isn't necessary (the data is too "fast moving" for there to be much point). This use case doesn't mean it's okay to lose data randomly. Cache invalidation should happen intentionally via an intelligent algorithm, and other data stores (such as Redis) provide this. > Here is another: A search engine for classifieds where the source of truth of 99%+ of listings were external feeds that'd get re-crawled at least once a day. If we lost a few million updates, who cared? If it was few enough, we'd just let the normal updates take care of it. If we had a major problem, or needed to do a major update that'd have compatibility implications, we could just do a re-crawl of all the feeds. Just because some data loss is acceptable doesn't mean it's desirable, and while I agree that most of the time some data loss doesn't matter, I don't think you can actually say it never will. What if the 1% listing happens to include an client's feed when the client is doing a major product launch? > You're scaling reads by replication, and so the vast majority of your data exists in multiple data stores. You may choose one you consider reliable for the source of truth, and the rest are basically caches, but you may want/need something more capable than a straight up key/value store for various reasons. Again, cache invalidation should happen intelligently via a well-thought-out algorithm, not by randomly dropping data, and there are solutions which supply that. > I believe the vast majority of data stores I've worked with have been ok to suffer a total loss of because most of them are not the source of truth, and the source of truth can be re-queries fast enough for it to not be a big deal to deal with losses in secondary copies. Okay, I don't buy this, but let's say it's true. What about the memory leaks? And what if your needs change, and you're no longer okay with data loss? Are you willing to take the risk that you're going to have to rewrite your storage layer because you chose a data store that drops data randomly, and your needs changed?
- marcosdumay 10y agoNo idea about what MongoDB would be good for, but: > In what case is it every acceptable for a data store to lose data? This is usually acceptable on analytics data, session data, on most of the "big data" applications... In fact, data that you can not lose whatever happens is the exception, not the rule. Anyway, when in doubt, it's certainly better to err to the side of security, not risk. Even more because all that security is already written into some middleware that you can simply install and use.