3 ms·
Tokenization was a simple string split on whitespace, and no stemming. It was quite a large Mongo dataset, so only a fraction of the index and data would've be
by ChrisFulstow 16y ago
Tokenization was a simple string split on whitespace, and no stemming. It was quite a large Mongo dataset, so only a fraction of the index and data would've been in memory, it could've easily been quicker for a smaller dataset living in memory. For me, one of the benefits of Lucene is the powerful built-in query parsing, tokenization, analyzers, etc.
- rabidsnail 16y agoThe performance probably would have been better if the dataset (or at least the portion of the dataset that gets touched frequently) was smaller. But why wouldn't lookup in lucene be at least as slow?