4 ms·
Lucene has the concept of "minimum match," [1] which is what you're playing with. You can implement your algorithm on top of this relatively trivially, by re-q
by fizx 5y ago
Lucene has the concept of "minimum match," [1] which is what you're playing with. You can implement your algorithm on top of this relatively trivially, by re-querying with min_match=n, min_match=n-1, etc until you get results. If you wrote a Lucene plugin to do this, it would also be small (<600 lines) and fast.
People have chosen not to do this because (1) it adds complexity, and (2) the bold declaration you made (A good top-K algorithm should rank a document containing more user query terms higher than a document containing less number of user query terms) turns out to be only somewhat true in practice. It's pretty common for pagerank and other non-linguistic ranking signals etc to trump linguistic preciseness.
That said, everything is a tradeoff, and the industry has clearly chosen to return more candidates than strictly necessary, then prune them after the fact with better ML & relevance, rather than trying to construct the perfectly minimal candidate set while missing some possible results.
[1] See https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-minimum-should-match.html https://www.elastic.co/guide/en/elasticsearch/reference/curr...
- huahaiy 5y agoThe industry has not "chosen" to not do what T-Wand does, the industry did not have that option because T-Wand just comes out. Now with T-Wand, the industry has that choice to make. It is easy to implement T-Wand in the existing Lucene code base, and I would hope so. Otherwise, it would be a shame.