8 ms·
Call me maybe: MongoDB (2013)
- ulisesrmzroche 12y agoThis is ancient stuff though, is this still relevant today?
- leif 12y agoYes, there are still problems with the election protocol, e.g. [1]. The right kind of network partitions can cause multiple primaries to stay up indefinitely, accepting writes on both sides of the partition, which will eventually be rolled back. There is another problem with the election protocol that allows writes acknowledged by a majority of machines to be rolled back after an election. Both of these problems can be fixed by using something like Raft[2] or Paxos for elections, rather than the ad hoc mechanisms used today. In TokuMX[3], we're currently working on replacing the election algorithm with something similar to Raft, that will eliminate these sources of data loss. We've heard that MongoDB is also working on fixing replication, but we don't know what their exact plans are (they have a bigger challenge since they need to stay compatible with their existing replication algorithms, which use timestamps as transaction identifiers) or whether these fixes will end up in 2.8 or in a later version. [1]: https://jira.mongodb.org/browse/SERVER-9848 https://jira.mongodb.org/browse/SERVER-9848 [2]: https://ramcloud.stanford.edu/wiki/download/attachments/11370504/raft.pdf https://ramcloud.stanford.edu/wiki/download/attachments/1137... [3]: http://docs.tokutek.com/tokumx http://docs.tokutek.com/tokumx
- ericingram 12y agoOnce again, a TokuMX engineer steps up to explain issue and offer a potential solution. I can't help but wonder why MongoDB engineers aren't doing this. But no matter, just glad we're using TokuMX.
- nevi-me 12y agoI think if I was working from someone's codebase, with about 80-90% of what I need already built-in, I would have the time and resources to make improvements on my fork, and shine glorious over the software making the base of my product. I wanted to try TokuMX months ago, but when I learnt that the version at the time was based on Mongo 2.2 I shied away from it, because I need GeoJSON capabilities. I remember that with 2.6 one of TokuTek's engineers said that they needed to look at Mongo's code and start playing catch-up, I don't know if they've done that so far. What will Mongo 2.8 mean for TokuMX? We're seeing document-level locking, possible B-Tree improvements (I presume Toku's R-Tree/Fractals [can't remember which they use] will still be superior), possible transactions (although what's on JIRA hasn't convinced me so far) and a few other improvements and Performance Boosting Things. So to what scale with Toku remain relevant if they don't keep up to date with Mongo, because in my case, using their versions based on 2.2, their ideology of being 'a drop-in replacement for MongoDB' doesn't work. I'll go to their Github page and try see whether they've merged the 2.6 codebase to their latest versions though :) EDIT: from looking at their release changelogs, as of October last year, they were in parity with Mongo 2.4, with the exception of geo-indices and full-text search, and 2.6 is still an open milestone. It kind of feels like the Joyent vs Strongloop thing on Node.js, but I wonder if TokuTek employees push bug-fixes upstream to Mongo, or whether they just fix them on TokuMX and use that as a selling-point; again with this I'll have to do some digging to inform my opinion, but I'd appreciate if someone who knows could clarify it.
- zardosht 12y agoAnother engineer at Tokutek here. As you see, we are up to 2.4, and have been investigating 2.6 and Geo. With all possible features, whether they be from MongoDB 2.6 or things we innovate on our own like partitioned collections, we prioritize and address them based on customer and user feedback. Also, 2.6 is not an all or nothing proposition that needs to be done in one release. Features with the most demand (whether it be the new write commands or aggregation framework improvements) will be done before others. We've done this before. When we released 1.0 that was based on 2.2, we also released hash based sharding with it which was a 2.4 feature. We did so because users demanded it. As for pushing bug fixes upstream, we file bugs when we see them. Our VP of engineering was a winner in the MongoDB 2.6 bug hunt with SERVER-12878. SERVER-9848 and SERVER-14382 are among the bugs I've filed.
- deleted 12y ago[deleted]
- onedev 12y agoPeople still use MongoDB in 2014?
- aioprisan 12y agoreally, what's the point of that comment?
- adamors 12y agoFor what it's worth, I think it an interesting question. Considering just how hyped Mongo was a few years ago, and how many other NoSQL databases (CouchDB, Riak, Cassandra perhaps even Postgres with its JSON support) gained mind share since then.
- onedev 12y agoYeah Mongo was "The Flavor Du Jour" just a couple years ago, with MongoDB articles hitting frontage like every week. Now we don't ever hear about it and people don't really use it much either.
- lwhalen 12y agoPeople don't use it because it is terrible. I had to manage a 5-node Mongo cluster as recently as 2 years ago, and I still drink heavily because of it. It was a data roach-motel - your documents check in, but they never check out!!! Everything ran fine, until it was time to fail over - and the company had a 'full failover every quarter' commitment. Every single time we went from the 2-node primary site to the 3-node secondary, all he would break loose. Bad writes, lost partitions, it was a new 'thing' every quarter. Mongo, in my limited experience, was a terrible platform and I recommend everyone who is thinking of implementing it to run screaming. It accepts data SUPER fast, but so does /dev/null.
- nevi-me 12y agoNot advocating MongoDB, but I think if we base our opinions solely on how things were 2 years ago, we would end up still basing our opinions on how things were 10 years ago. Software breaks, 'poor' design decisions are made, etc, but at the same time, 2 years is a lot of time for improvements to be made. My HTML5 app failed me 3 years ago, I've now been recommending that everyone stay away from HTML5. /s
- olegp 12y agoFor all the hate MongoDB gets, I have to say building queries in server side JS using MongoDB JSON syntax rather than SQL style queries is the way to go. Check out these two examples: https://github.com/olegp/stick-blog-pg/blob/master/lib/server.js https://github.com/olegp/stick-blog-pg/blob/master/lib/serve... - uses Postgres https://github.com/olegp/stick-blog/blob/master/lib/server.js https://github.com/olegp/stick-blog/blob/master/lib/server.j... - uses MongoDB - much easier to construct dynamic queries using the JSON syntax There are of course plenty of other things a relational DB like Postgres has going for it, so I've been experimenting with having a MongoDB interface to Postgres with data stored using the JSON datatype. It's now feature complete and passes all unit tests with read performance exceeding that of Mongo in some benchmarks: https://github.com/olegp/pg-mongo https://github.com/olegp/pg-mongo
- goldenkey 12y agoYour comparison is not apt; an ORM can be used for any SQL flavor.
- olegp 12y agoHaving written an ORM myself back in 2000 and used Hibernate and JPA for many years, I'm now of the opinion that ORMs are an unnecessary, leaky abstraction in the modern day of dynamic languages. One is better off talking directly to the data store and effectively streaming data out via a REST API.
- rspeer 12y agoWhy would someone downvote this? ORMs certainly are a leaky abstraction, and they certainly do come with a penalty in performance and expressiveness. Is it really downvote-worthy to add the opinion that they are unnecessary?
- deleted 12y ago[deleted]
- 12y ago
- Gonzih 12y agoWith rise of RethinkDB it would be lovely to see similar post on it.
- hopeless 12y agoI've looked into RethinkDB and I really like it but… you have to migrate your data between each 1.x -> 1.y release which might be trivial early on but impossible at a larger scale :-/ http://rethinkdb.com/stability/ http://rethinkdb.com/stability/
- neumino 12y agoThis will not be needed for the next releases (if everything goes as planned) https://github.com/rethinkdb/rethinkdb/issues/1010#issuecomment-47996409 https://github.com/rethinkdb/rethinkdb/issues/1010#issuecomm...
- garretraziel 12y agoI'm building small application using Node.js and MongoDB and I'm planning to host it on Openshift or Heroku. All that hate that MongoDB takes here on HN makes me reconsider technologies I am using. I will not have many relations in my database (model User, model Document, User owns Document... and that's all) so I thought that NoSQL databases will do. Plus, MongoDB lets me use GridFS - I'm planning to store pdf presentations in it. If I should drop MongoDB, what other technology should I use? Or should I fall back to Postgres + ORM and manage my files in filesystem manually? I don't want to start a flame, I am looking for an advice. I have considered MongoDB to be "good enough" as GridFS lets me store my files without a hassle, but after all that I read on the Internet, now I am not so sure.
- barkingcat 12y agoI don't think it's hate - it's a call for users to understand and examine the implications of assumptions that the developers of the tools have made. Don't take the advertising taglines and slogans as a panacea - there are still a ton of things to understand about how mongodb might or might not work for your circumstances. The advertising buzzwords are just that - advertising. You need to examine each and every technology that you chose as part of your project. And I totally recommend the entire Jepsen series from that site http://aphyr.com/tags/Jepsen http://aphyr.com/tags/Jepsen - it's an eyeopener! In my opinion, if you are starting some experiments, etc, just use postgres! It removes a lot of doubt and new learning/confusion from your development.
- Rapzid 12y agoPostgreSQL actually works quite well as a noSQl/document store :) I believe the best answer though is "it depends". I'm not sure this is the best place but if you gathered your requirements and listed them I'm sure a dstabse that will keep your data safe can be recommended.
- reqres 12y agoI was in a similar situation to you when MongoDB was a lot less mature. Postgres was my "go to" store and I was a very skeptical of the Mongo hype. My view would be to use Mongo if it's just a small project. It's a good time to learn exactly how Mongo is different and get a first hand experience of what the trade-offs are.
- StavrosK 12y agoIs there any datastore in this series that behaved correctly in a partition? I've seen ElasticSearch, Redis, Riak, Mongo, and all of them crapped their pants.
- deleted 12y ago[deleted]
- misframer 12y agoZookeeper? Not in the same category as MongoDB, of course.
- nemothekid 12y agoCassandra, Zookeeper and Kafka didn't shit the bed.
- tveita 12y agoKafka did lose a lot of data in the Jepsen test, as it will by design to preserve availability under some specific circumstances. The behavior in question will be configurable from Kafka 0.8.2, by setting "unclean.leader.election.enable" to false. See https://issues.apache.org/jira/browse/KAFKA-1028 https://issues.apache.org/jira/browse/KAFKA-1028
- dbenhur 12y agoOn Cassandra, some CQL collection operations behaved well (adding elements to a set). Everything else Kyle tested there was demonstrated to lose data. Kafka lost data all over the place under Jepson; though Kyle offered great respect to the team and expects them to deliver optionally configured safer semantics (at a performance cost) in future releases. Riak was totally solid when using CRDTs and turning off the insane LWW default.
- zardosht 12y agoAnother thing to keep in mind is that not all systems are meant to be CP systems.
- brickcap 12y ago
- StavrosK 12y agoIs there any datastore in this series that behaved correctly in a partition? I've seen ElasticSearch, Redis, Riak, Mongo, and all of them crapped their pants.
- _JamesA_ 12y agoDoes anyone have experience/comparisons with OrientDB? http://www.orientechnologies.com/orientdb/ http://www.orientechnologies.com/orientdb/ It seems to be much lesser known but it ticks all the right boxes compared to other NoSQL datastores.