4 ms·
I'm an engineer at Tokutek I'm confused by your comment. The beginning acknowledges the fact that MongoDB has a weak storage engine, but your conclusion is th
by leif 13y ago
I'm an engineer at Tokutek
I'm confused by your comment. The beginning acknowledges the fact that MongoDB has a weak storage engine, but your conclusion is that, even with a strong storage engine like ours, there is still a problem. What other problems do you see? Are they something we could work on?
- jpgvm 13y agoThis is going to come off as abit negative but I kinda feel it has to be said. I would first like to say I do love the Fractal tree indexing, very cool and could have alot more intesting usecases outside of databases (I'm thinking logical volume/block storage etc.. I'm always thinking in kernel land..) The problem is that Mongo advertised itself as a database and wasn't one. Once you do that reputation of the product is dead forever. TokuMX is a real database as far as I can see, MVCC, great indexing story etc. By association TokuMX is probably not regarded as highly as it should be. Which is a shame but it's a people problem, not a technical one. People can very easily lose trust in a technology at which point it's effectively dead, it might take a long time to die due to lock-in but it's dead. For instance I have recently started playing with RethinkDB over TokuMX almost purely because of Mongo association. Now technically that might not sound like good reasoning but when you think about the kind of person that writes a database that doesn't fsync your writes by default and relies on the page-cache over doing direct I/O when building a database.. doesn't really inspire confidence in the network stack, the query planner.. or well anything. If anything it makes you insistent on not having ANYTHING to do with that sort of codebase. Just replacing the storage engine might actually be good enough, but restoring my trust in the rest of the codebase is almost a forgone conclusion at this point.
- leif 13y agoI've seen a lot of the rest of their code, and most if it is getting better over time, as they grow they're forced to adopt better habits in order to scale their engineering team. I think you're misunderstanding the type of programmers they are. They didn't use mmap because they are sloppy everywhere, they used mmap because their critical innovation was not in storage. What they really thought was valuable, what they wanted to work on, was the query language and cluster management tools, so they did the simplest thing for storage and moved on (personally I don't understand why they didn't just use BDB, maybe they were afraid of transactions, but I suppose everyone has a little NIH syndrome in their database). Now they're a bit locked in to that code, because after bolting on journaling (that architecture is a brilliant but incredibly dirty hack), the code is a mess and I'm sure nobody wants to touch it. In fact most of the other subsystems have been getting cleaner rewrites, except for the storage layer. I think the only way out is a complete replacement, which is what we did so I feel pretty good about that. So I don't know if I'll convince you, but I've read a lot of their code (especially in the last few weeks, I've been backporting things from 2.4), and that's the feeling I get about their history and vision. Hope it gives you some insight.
- chris_wot 13y agoThis post does not inspire confidence. Sorry, has to be said.
- willvarfar 13y agoAs I recall it, all the early noise they generated was their excitement about their hot benchmarks and how good mmap was... They can try and rewrite the web and remove all the silly benchmarks, but they were the loudest "web scale" cowboys back in the beginning and we remember them for it.
- jpgvm 13y agoI don't think the problem is I misunderstand them, I just disagree with them I disagree with them on what is the minimum viable product for a database. I come a storage and service provider background where failures are treated very harshly (usually death of companies for singular mistakes) so I take releasing a product that stores customer data very seriously. To be honest this is the biggest attraction for me to RethinkDB. They waited a sufficiently long amount of time with a commercially backed team of very competent engineers that obviously have the required background to sit down and DESIGN a database. The query language generates a non-turing complete language with a clean AST the has all the right deterministic characteristics to implement a powerful planner/optimizer. Their on disk format has been abit in flux but the core design is excellent and you can see that it has been optimized for very fast range queries. Even the API protocol and serialization were designed with care, not to mention the excellent ReQL language and attention to detail when integrating drivers into the host language. Which is the other thing I tend to dislike about Mongo, it reeks of lack of design. The journalling effort for instance as you pointed out is very adhoc, this goes for GridFS and alot of the other features they have integrated into the codebase. These are smells that I can't ignore when looking at a product that I need to trust with my data. The counter argument is to not trust it with your data. But I am yet to find a reason where that makes sense where another datastore wouldn't be a better choice.