4 ms·
This isn't really supposed to be a proper full-text index feature, instead it's building a rudimentary inverted index using a string array property. It's possi
by ChrisFulstow 16y ago
This isn't really supposed to be a proper full-text index feature, instead it's building a rudimentary inverted index using a string array property. It's possible to create indexes over array properties in MongoDB, which is very cool, and increases performance to an extent. But even with an index, this approach to full-text was much slower for me than an equivalent search against the same data in Lucene.
I'd love to see a MongoDB component that replicates data from the oplog to a dedicated full-text store like Lucene or Solr.
- rabidsnail 16y agoHow are you doing your tokenization and stemming? I find it hard to believe that the actual token lookup is slower.
- ChrisFulstow 16y agoTokenization was a simple string split on whitespace, and no stemming. It was quite a large Mongo dataset, so only a fraction of the index and data would've been in memory, it could've easily been quicker for a smaller dataset living in memory. For me, one of the benefits of Lucene is the powerful built-in query parsing, tokenization, analyzers, etc.
- rabidsnail 16y agoThe performance probably would have been better if the dataset (or at least the portion of the dataset that gets touched frequently) was smaller. But why wouldn't lookup in lucene be at least as slow?
- rb2k_ 16y ago> I'd love to see a MongoDB component that replicates data from the oplog to a dedicated full-text store like Lucene or Solr. Photovoltaic does this: https://github.com/mikejs/photovoltaic https://github.com/mikejs/photovoltaic I haven't had time to play arround with it yet though :(