3 ms·
Lucene, and most other indexing solutions (to my knowledge) index based on words, not arbitrary substrings. So they can efficiently answer questions like "Give
by nelhage 14y ago
Lucene, and most other indexing solutions (to my knowledge) index based on words, not arbitrary substrings.
So they can efficiently answer questions like "Give me all documents containing the word 'Linus'", but not necessarily "All documents containing the string '#def'".
That said, I am indexing -- I have my own custom backend that stores an in-memory index that lets me do arbitrary substring search (and more complicated queries, such as most character classes) much faster than a full search.
That said, the backend will fall back on a full regex search if necessary.
- durin42 14y agoDo you use anything like a trigram index (see rsc's wonderful posts about how Google Code Search worked, and https://code.google.com/p/codesearch/ https://code.google.com/p/codesearch/ for a Go implementation) to speed up the regex codepath?
- nelhage 14y agoI'm using a different data structure -- a suffix array -- but the concept is pretty similar. I started work on this before Russ released his codesearch implementation, but I did read his blog posts while I was working on this.