4 ms·
Building upon @m_ke's note, when you have the sentence embeddings indexed, you use query similarity with those sentences as an additional signal that can be use
by binarymax 6y ago
Building upon @m_ke's note, when you have the sentence embeddings indexed, you use query similarity with those sentences as an additional signal that can be used further. For example, you can weight sentences that appear earlier in the document higher, you can take the top N sentences and use those, you can apply clustering to find hotspots, etc.
I'd treat it as a base signal like BM25 for a field - which is frequently tuned and used as one of many features for relevance.
- petulla 6y agoSort of makes sense. You'd need a ground-truth search->document dataset, I think, to test on, though. Seems a bit hand-tuned.
- binarymax 6y agoAbsolutely! We call it a "judgement list" in search relevance land, and one is required and used for training and testing of a relevance config/model...though the training is often manual, until your configuration and infrastructure are mature enough for learning-to-rank.
- petulla 6y agoLast q: which datasets of this kind are used for SOTA testing or document clustering?
- binarymax 6y ago@petulla we've reached the thread depth limit so I can't reply to your last question, but typically MSMarco is used for SOTA testing of deep learning for search: https://microsoft.github.io/msmarco/ https://microsoft.github.io/msmarco/