7 ms·
MongoDB's Write Lock
- milkshakes 15y agothank you for this. i don't understand why 10gen didn't put something out like this in the first place, it would definitely helpfully frame a lot of the more annoying discussions i've had.
- rick446 15y agoGlad to help! 10gen actually has a policy of never sharing benchmarks so that explains why they never said anything.
- xxqs 15y agoWhy is everyone paying so much attention to MongoDB? It has been criticized a lot for its design and implementation problems, but still for some reason it's so popular. To name a few, * word-unaligned memory structures, which leads to incompatibility with virtually any non-x86 CPU architecture * explicitly little-endian processing in the server, so there is no way to run the original code on any big-endian CPU architecture. There has been a patchset which was tested on a SPARC CPU, but last time I asked the author, the 10gen team completely ignored this effort. apart from that, there have been reports of data loss without any failure note
- fleitz 15y agoDidn't you hear? Mongodb is webscale and runs in the cloud. That's all you really need to know, please ignore any rational arguments and just repeat webscale and cloud endlessly. Remember that if you run into scalability issues in the cloud all you have to do is spin up 300 instances to get the performance of one 5400 rpm laptop drive. The cloud is webscale your laptop is not. Also please ignore that a laptop with an ssd will need about 3000 instances to get the same performance.
- ketralnis 15y agoYeah, how dare people like what you don't like? > It has been criticized a lot [...] but still for some reason it's so popular Since you provide no data or sources for "criticized a lot" it's no surprise that you don't provide the same for "so popular". I assume you mean "I've seen some headlines on Hacker News about it". > incompatibility with virtually any non-x86 CPU architecture [...] no way to run the original code on any big-endian CPU architecture Huh. Maybe that's only a problem for people on non-x86 CPUs then? Look I'm really not a fan of Mongo but "It has been criticized a lot [...] but still for some reason it's so popular" describes every technology ever. Get over it.
- xxqs 15y ago> Maybe that's only a problem for people on non-x86 CPUs then? actually it's a problem for application developers. I cannot rely on a backend system with limitations like these. So I'll have to go back to the RDBMS backend or look for other nosql alternatives, but definitely mongoDB is off my list
- nknight 15y agoThe world runs on x86. You might have some legacy systems running SPARC or POWER, but those systems are unlikely to reside in MongoDB's target market anyway. Arguing everyone else should ditch code that happens to not work on your pet architecture is a pretty self-centered worldview.
- xxqs 15y agoactually the number of ARM processors is growing, and not only in the mobile sector. There have been some efforts to bring ARM architecture into the server market. Also China is building its own MIPS-based supercomputer. Also the SPARC architecture is actually developing, although it's a pity to see it swallowed by Oracle. IBM is still shipping PowerPC servers. besides, there are huge SPARC-only datacenters still running.
- nknight 15y ago
- nknight 15y ago> In MongoDB version 2.0 and higher, this is addressed by detecting the likelihood of a page fault and releasing the lock before faulting. I'm assuming MongoDB tries to detect this with OS-specific syscalls. Has there been any attempt to determine whether it would be even faster and/or more portable to just unconditionally "read" the pages before acquiring the lock?
- latch 15y agoI forget who, but a fairly popular implementation of MongoDB once posted about their experience, and they mentioned that they always did a find before doing an update. Every now and again you'll see this approach get suggested in the groups.
- dolinsky 15y agoSounds like a great extension of the benchmarks provided in the article.
- fleitz 15y agoWhy not just go take the write lock out of the db all together if it doesn't need to be journalled? It's obvious at that point that missing / mangled data is acceptable. It's not particularly amazing that not writing data to disk is faster than writing to disk. What IS amazing is that the geniuses at 10gen have some how managed to make not writing to disk thousands of times slower than writing to disk. Who would design a database so shitty that journalling the data impacts performance. Typically you only need one spindle for a journal to support 100 to 200 data spindles. If you can't pull 80 to 90 mb/sec sustained write from a log drive something is seriously wrong. 48 iops now that's what I call "web scale". Let me just throw out my acid database that does 30000 iops on Win2k3 of all things, to get 48 iops with out journalling.
- peschkaj 15y agoMongoDB's journal is another collection that only syncs to disk 10 times a second. It's not a true journal like a write-ahead log. You can force that collection to fsync, but... yeah, I get get stupendous performance using a cheap RAID controller on commodity hardware with an old and busted relational database. Hooray for 40 years of technology!
- DenisM 15y ago> MongoDB, as some of you may know, has a process-wide write lock. I've never taken time to see what MogoDB db is, but thanks to this opening sentence, now I know everything I ever wanted to know about this system. Having worked for 13 years on database system design I am pretty confident that a system not designed with concurrency in mind cannot be retrofitted with any decent concurrency later. Thank you Rick for saving me the time.
- cperciva 15y agoI'm not convinced. I'm not saying that MongoDB is well designed -- I don't think it is -- but it seems to me that a process-wide write lock would be perfectly fine for a data store which is designed to cluster at a one-process-per-CPU-core level.
- DenisM 15y agoSo one core is processing one request at a time, right? If it spends any time at all being queued up on disk IO or network traffic or anything else, that core is burning up XX watts for no good reason. A more efficient system, designed for concurrency, will have higher HW utilization and lower cost. For a sign of things to come, I invite you to take a look at how relational database vendors are fighting to squeeze single-digit percentages from their engines to beat benchmarks such as TPC-C or TPC-E. NoSQL will get there too - fight for efficiency.
- nknight 15y ago> queued up on disk IO And there's where your legacy understanding fails. You're really not supposed to be doing a lot of disk IO in modern datastores in the first place. Your working set should be in RAM. > beat benchmarks If my vendor is investing time in beating benchmarks instead of solving real problems, I'm finding a new vendor.
- eternalban 15y ago"solving real problems" I am not at all addressing MongoDB here to be clear -- just your comment regarding worthwhile "problems" to "solve". It is not too difficult to foresee a future where energy costs will trump all other considerations, including development time, sys ops, etc. Specially for <X>aaS efforts, energy efficiency can clearly end up being a competitive edge.
- ilaksh 15y agoThe thing is that you are still supposed to keep the whole working set in memory and use sharding if its larger than that. Which means that none of this is really relevant. Except that now with that new graph showing such good performance on reads during paging people are going to get confused. Anyway you can get 32GB of RAM for $232 or 48GB for $636. Which means that for 90% of applications, you actually don't need to shard. And if the journaling works (people with a vested interest in relational systems really have to hope that it doesn't), then I don't have to worry about my data disappearing, even if I only have one database server, which is also not a recommended design with MongoDB (or any database really). I have years of experience with SQL Server, Oracle and MySQL. However, MongoDB is the most attractive database now because it makes the object-relational impedance mismatch go away. http://en.wikipedia.org/wiki/Object-relational_impedance_mismatch http://en.wikipedia.org/wiki/Object-relational_impedance_mis... If I can write code (in CoffeeScript using a library like Mongolian Deadbeef) like this posts.insert pageId: "hallo" title: "Hallo" body: "Welcome to my new blog!" created: new Date posts.findOne pageId: "hallo" , (err, post) -> posts.find().limit(5).sort(created: 1).toArray (err, array) -> then whey would I want to deal with separate steps of setting up the relational database tables, creating stored procedures, creating a software layer to map my objects to my tables etc., or hiring a DBA? I believe that most of the hate for MongoDB is fueled by a survival instinct. The popularity of databases like MongoDB threaten to make years of experience obsolete and threaten the existence of the DBA profession. Relational databases are great, but they were an optimization designed to solve certain problems that most people today just don't have, and now they have become an unfortunate institutionalized dogma.
- JanezStupar 15y agoFrom multiple years of NoSQL experience. The Object/Relational impedance mismatch stays right there where you leave it. Using K/V stores will only help you write data. But the whole impedance nightmare is right there, waiting for you to try and make some sense of the data. Especially if you want to do relations. And you are dead wrong about relational DB's. It is either due to habit or because of being a better fit that users demand representations of data that are best served from a relational source. So you better plan your data models wisely, because you WILL pump this data into a relational source, sooner or later. It would be wise do keep a schema around all the time. In the end it is merely a question of data normalization and the use case at hand.
- amalag 15y agoI think Mongo is great for some use cases. There are some use cases where the flexible json data just makes sense. Regarding his benchmarks, he turned off journaling. Would love to see them with journaling turned on, see how much is relevant.
- steve8918 15y agoDoes anyone know what the performance differences would be between MongoDB and SQL Server/Oracle if they all had enough RAM to hold the entire dataset in memory? I'm only guessing but it seems to me that any database with their entire dataset in memory would be very fast, no?
- peschkaj 15y agoYou are correct - any database with the entire dataset in memory will be incredibly fast. SQL Server bypasses the Windows file system cache and will aggressively manage memory to keep frequently accessed data pages resident in RAM. The read/write performance is what you would expect for a database with fine-graned lock management - when you have to go to disk things get slower, otherwise I/O is only limited by RAM and the overhead of lock management.
- megaman821 15y agoThere were some slides showing PostgreSQL with fsync turned off performed about the same as MongoDB. There is no MongoDB secret sauce that makes it any faster than well established relational dbs with a few configuration tweaks to make the comparison even.
- orthecreedence 15y ago> If you are able to do this, it turns out that the global write lock really doesn't affect you. Blocking reads for a few nanoseconds while a write completes turns out to be a non-issue. (I have not measured this, but I suspect that the acquisition of the global write lock takes significantly longer than the actual write.) Actually, it does affect you. I have worked with mongodb in production in a high-write scenario with about 1000 clients and it slowed to a crawl. All data was in memory. The server was not breaking much of a sweat. mongostat showed upwards of 60 queued reads/writes at any given time. The only solution was to shard, but I feel like an enormous server like the one we were using should be able to handle 1000 writing clients. Keep in mind this was version 1.8. I no longer work at the company where this happened and cannot testify to the performance of 2.0, but 1.8 has abysmal write performance.
- mnutt 15y agoWere your writes changing the size of the documents so that mongo had to move them? I've had this happen and it'll cause mongo to grind to a halt.
- orthecreedence 15y agoSome of them were, but others were just updating a boolean or an integer value. For the most part we tried to pad our records, but I'm sure there was some moving along the way.
- rick446 15y agoI guess it depends on what % of your writes were simply updating a boolean or integer value. My benchmark shows that simple updates like that don't affect query performance much. Writes that take longer probably have different performance characteristics, YMMV, etc.