10 ms·
MongoDB 2.4 Released: Text Search, Security, Hash-based Sharding
- dkhenry 14y agoI am curious to play with the text searching. I wonder how it stacks up to Lucerne or Solr in the text indexing space
- mumrah 14y agoI am quite sure it will not come close to Lucene/Solr in terms of performance or capabilities.
- andersnolsen 14y agoOnly one thing to do - test
- Argorak 14y agoOr just read the blog post: tl;dr: use it, if a simple reverse index fits your needs, use Lucene for the grown-up stuff > MongoDB text search is still in its infancy and we encourage you to try it out on your datasets. Many applications use both MongoDB and Solr/Lucene, but realize that there is still a feature gap. For some applications, the basic text search that we are introducing may be sufficient. As you get to know text search, you can determine when MongoDB has crossed the threshold for what you need.
- ecaron 14y agoThe MongoDB staff had said they're mostly adding text searching to satisfy a repeat request that they've heard over and over - but they've said that anyone really wanting to do text searching at scale should use something like Sphinx, Solr or ElasticSearch.
- tbrock 14y agoKeep in mind that it is both new and experimental. If you are comparing them based on functionality and performance you are missing the point. It is my opinion that this easily bests external implementations by reducing moving parts and removing the asynchronous nature of the updates to search indexes. The text index is updated atomically and in real time. Best of all: it's not another part of your stack that needs to be setup and maintained.
- deleted 14y ago[deleted]
- c-oreills 14y ago"db.killOp() Can Now Kill Foreground Index Builds" That's going to save a lot of accidental "shit-the-whole-db-is-locked" pain.
- vailripper 14y agoLove seeing the text indexing!
- dmytton 14y agoUsing the new working set analyser is going to make figuring out how much memory you need significantly easier. Giving MongoDB enough memory for your working set is the easiest way to get performance but it was quite difficult to figure it out: you had to know both data and index sizes across your most common usage patterns.
- diminoten 14y agoDo you know of any resources that give some info on this topic in general? I've been looking but I haven't found anything that's sufficiently plainly-spoken that I fully grasp the contents.
- francesca 14y agothis presentation offers a good overview pre 2.4, but will give you insights into how you can prepare for memory issues http://www.10gen.com/presentations/mongosv-2012/capacity-planning http://www.10gen.com/presentations/mongosv-2012/capacity-pla...
- kstirman 14y agofrom: http://docs.mongodb.org/manual/reference/server-status/ http://docs.mongodb.org/manual/reference/server-status/ serverStatus.workingSet.pagesInMemory pagesInMemory contains a count of the total number of pages accessed by mongod over the period displayed in overSeconds. The default page size is 4 kilobytes: to convert this value to the amount of data in memory multiply this value by 4 kilobytes. If your total working set is less than the size of physical memory, over time the value of pagesInMemory will reflect your data size. Conversely, if your data set is greater than the size of the physical RAM, this number will reflect the total size of physical RAM. Use pagesInMemory in conjunction with overSeconds to help estimate the actual size of the working set. Also, there's a discussion of working set in this guide starting on page 9: http://info.10gen.com/rs/10gen/images/10gen-MongoDB_Operations_Best_Practices.pdf http://info.10gen.com/rs/10gen/images/10gen-MongoDB_Operatio...
- badgar 14y ago> you had to know both data and index sizes across your most common usage patterns. This is called capacity planning. If you can't estimate what resources you use, I wouldn't expect great success scaling.
- leif 14y agoI'm glad this is out. I'm really going to enjoy plugging fractal trees in under FTS.
- davidkellis 14y agoI was hoping for collection-level locking to be a part of the 2.4 release. I didn't see any mention of it in the release notes. Last I heard they were going to implement collection-level locking and then begin work on document-level locking. I'm still hoping document-level locking isn't too far off.
- spf13 14y agoCollection level locking isn't in 2.4. Working on more granular locking for 2.6/2.8. May skip collection level entirely for something more granular. More details can be found in the collection level locking ticket https://jira.mongodb.org/browse/SERVER-1240 https://jira.mongodb.org/browse/SERVER-1240
- davidkellis 14y agoSkipping collection level locking and going for something more granular (like document level) would be awesome.
- tomsthumb 14y agoDon't you pretty much get document level locking by using multiple $<update> modifiers in a single query?
- apendleton 14y agoFrom an atomicity perspective, yes. From a performance perspective, no; other concurrent operations affecting anything else on the whole database wait on your write.
- c-oreills 14y agoYou can circumvent this in the short term by using one collection per db.
- camus 14y ago
- leothekim 14y agoAccording to the upgrade instructions [1], the only supported upgrade path from sharded 2.0 clusters is via 2.2. [1] http://docs.mongodb.org/manual/release-notes/2.4-upgrade/#upgrade-a-sharded-cluster-from-mongodb-2-2-to-mongodb-2-4 http://docs.mongodb.org/manual/release-notes/2.4-upgrade/#up...
- c-oreills 14y agoAnd infact you can't have a mix of 2.2 and 2.4 boxes in the same cluster, you have to go via 2.2.1
- gregjor 14y agoGreat to see the Pick database reinvented bit by bit. Takes me back to 1980. Forget the past, doomed to repeat, etc.
- julien_c 14y agoCan you elaborate for those of use who weren't around in 1980?
- tomsthumb 14y agoOr 1990.....
- shin_lao 14y agoBasically the guys from MongoDB are rebuilding databases as they existed in the 1980s with the motto "we will do better than relational databases!" or "Sybase, Oracle, prepare to die!". It's not clear which problem MongoDB is trying to solve or if it is an improvement over existing technology (this is my personal opinion).
- base698 14y agoI think it's pretty clear to anyone who's used it. It allows for rapid, low overhead changes to your data model in the very early stages of building something new. Of course, the further along you get those changes are no longer low overhead, but at the start of the new project it's very easy to get up and running.
- gtaylor 14y agoI guess I just don't find schema that big a deal even in the early goings. A relational DB schema isn't necessarily hard to change and adapt on the fly. I don't feel like this is the best thing for Mongo to hang its hat on, and I'm not sure they really are. I think its users tend to point at this as a big feature more than they should. I feel like the other reasons to use Mongo should show up more in these discussions than this "It's easy to make schema changes". Perhaps these are easier scaling/sharding, I don't know, I don't use Mongo. I just see the schema part of this as a sidepoint (albeit an important one from an application design perspective). There are legitimate cases where supremely simple schema flexibility is desirable, but I suspect a lot of people who think they need this really don't (there are easy ways to do this with relational DBs), and are in fact making some things more complicated as a result. While it IS very easy to change your schema with Mongo, let's not forget that you're going to need to enforce certain schema-related rules in your codebase instead of the DB now.
- Lionga 14y agoIt is great to have a stable version with fast count indexes see https://jira.mongodb.org/browse/SERVER-1752 https://jira.mongodb.org/browse/SERVER-1752 Best feature for me
- friendly_chap 14y agoThey are not fast, they are "normal speed" now. They were slow before. That was one of the problems which made me question the competency of the Mongo team. Best feature for me too anyway.
- derricki 14y agoI was hoping the security enhancements would include SSL certificate validation. Anyone know why they don't do that, or how a user should approach that limitation?
- matthewlucid 14y agoWe'll be moving away from MongoDB because it doesn't support certificate validation. What is the point of SSL connections if you don't validate the certificate? It seems that you get all the drawbacks of encryption (overhead, throughput) with none of the benefits (security). I'd love to see a solution.
- milkie 14y agoSSL certificate validation is in the 2.4 release.
- mrinterweb 14y agoI'm guessing this is the ticket for SSL support https://jira.mongodb.org/browse/SERVER-7202 https://jira.mongodb.org/browse/SERVER-7202
- jonesjim 14y agoWhoop! Multithreaded javascript with V8 JavaScript engine!
- apendleton 14y agoI can't actually find any details about how this works in practice. Are multiple maps and reduces in a map-reduce executed simultaneously within a single mongod process? If so, how many? Is it based on the number of cores in the server? Edit: not asking you specifically, just generally curious.
- tbrock 14y agoEach map reduce job can still only use one thread. Before, under spider monkey, only one job could be executed per mongod instance. Now you can have many jobs executing in parallel on a single server instance but each one of them still only uses a single core.
- apendleton 14y agoDo you know if this is likely to be permanent? In an application I work on, I do map-reduces on Mongo data using a third-party framework at the moment specifically to work around the inability to easily parallelize map-reduce jobs on a single machine.
- ranman 14y agoV8 is another important change that isn't really touted much in this release.
- tbrock 14y agoYeah, this is a big deal. It enables very interesting opportunities to create a map reduce engine that isn't awful or which requires Java. People use Hadoop because they have to, not because it is great. Isn't it time for a better alternative now that the toothpaste is out of the tube re: map/reduce? Java doesn't even have hash literal syntax! Why would you ever want to query document oriented data with it as your language of expression?
- lucian1900 14y agoIt's not hard to write Hadoop map/reduce queries in other languages: mrjob for python is particularly nice. It's also not relevant that SpiderMonkey is being replaced with V8. Mongo's map/reduce is also just a toy, not at all comparable with Hadoop's and being deprecated in favour of the aggregation framework. Also, there already are decent alternative map/reduce implementations. Disco is a good example, with a similar design to Hadoop.
- kstirman 14y agoMongoDB's MapReduce implementation is not being deprecated. The primary beneficiary of the V8 implementation is MapReduce, so it should be seen as further investment in this area. You can also run Hadoop MapReduce over data in MongoDB directly (24 - 29): http://www.slideshare.net/spf13/mongodb-and-hadoop http://www.slideshare.net/spf13/mongodb-and-hadoop
- L0j1k 14y ago[Comment about supporting only one master]
- dschiptsov 14y agoany row-level locking?)