3 ms·
I have some experience in developing IR/search software ( https://github.com/phaistos-networks/Trinity https://github.com/phaistos-networks/Trinity ) and, the w
by markpapadakis 9y ago
I have some experience in developing IR/search software ( https://github.com/phaistos-networks/Trinity https://github.com/phaistos-networks/Trinity ) and, the way I see it, it all comes down to accepting the following premises:
- It’s all about the relevance models. Specifically, BM25 and TF-IDF which are widely used by e.g Lucene and variants just won’t do. Specifically, they only word for large enough documents anyway.
- Indexing and Search algorithms and practices haven’t changed much in decades (though some novel ideas have been introduced not long ago). Lucene’s index encoding is compact and facilitates fast access, but even compared to the one made available by Google which is arguable simpler in design, doesn’t result in more than around 5% reduction in index size and postings list access time (according to my measurements that is). Posting lists intersections, unions and other such operations implementations are pretty much common across IR systems as well, with little room for improvement.
To get great results, you need really great relevance models(1), a great query rewrite system(2), and, because query rewrites usually expand a query to include multiple disjunctions (OR terms and phrases), your search engine needs to be particularly efficient at handling those(3).
You also need to care for spelling suggestions and personalisation/content biases and factors, but those are secondary concerns.