2 ms·
I evaluated MongoDB as a search engine but what I didn't like was high overhead of using indexes for words. It was like 10x slower when inserting documents cont
by aartur 14y ago
I evaluated MongoDB as a search engine but what I didn't like was high overhead of using indexes for words. It was like 10x slower when inserting documents containing indexed string-array field. I got real-time indexes updating, but I wanted to have some kind of bulk index updates, as the overhead was too big. And it wasn't available when I checked it about 1 year ago.
I don't know what the plan is, but an integrated stemmer alone isn't of big value to me. Actually I prefer to have a stemmer in my own code, so I can tweak/update it without touching a database.
- saurik 14y agoYeah... I'm a big fan of PostgreSQL's included full-text search features, but it always seemed weird to me that they seemed to go to so much trouble to include a complex multi-language stemmer... I guess it made it easy to get started with, but I immediately found a bunch of application-specific "no, I really need these specific words (maybe the name of the company, or commonly-searched-for API function names that are all fairly similar) to not be stemmed, and I want you to index this special syntax (maybe a Twitter @-reply or hashtag) as a single token including the punctuation" issues that made me start providing pre-stemmed streams of tokens.