12 ms·
Good question. Averaging can work for short documents, but this becomes less useful as the document text gets longer. When dealing with documents of multiple
by binarymax 6y ago
Good question. Averaging can work for short documents, but this becomes less useful as the document text gets longer. When dealing with documents of multiple paragraphs, I've been gravitating more towards thinking "whats the most relevant part of the document?". And for this SBERT works very well - as you can pluck the document's most similar sentence(s) for the given query, and use that as a signal for the most relevant piece of information.
Of course there are many caveats here, and this is by no means a solved problem, but works well for queries that are more than a couple terms.
- petulla 6y agoI'm not sure I understand in practice how that would work. You have sentence embeddings and a query embedding. You're comparing groups of sentences from a document to the query. But how do you decide how to group the sentences? Are you just up-front grouping similar sentences somehow?
- m_ke 6y agoYou index an embedding for each sentence and match against sentence embeddings, then rank documents based on highest scoring sentence matches.
- binarymax 6y agoBuilding upon @m_ke's note, when you have the sentence embeddings indexed, you use query similarity with those sentences as an additional signal that can be used further. For example, you can weight sentences that appear earlier in the document higher, you can take the top N sentences and use those, you can apply clustering to find hotspots, etc. I'd treat it as a base signal like BM25 for a field - which is frequently tuned and used as one of many features for relevance.
- petulla 6y agoSort of makes sense. You'd need a ground-truth search->document dataset, I think, to test on, though. Seems a bit hand-tuned.
- binarymax 6y agoAbsolutely! We call it a "judgement list" in search relevance land, and one is required and used for training and testing of a relevance config/model...though the training is often manual, until your configuration and infrastructure are mature enough for learning-to-rank.
- petulla 6y agoLast q: which datasets of this kind are used for SOTA testing or document clustering?
- binarymax 6y ago@petulla we've reached the thread depth limit so I can't reply to your last question, but typically MSMarco is used for SOTA testing of deep learning for search: https://microsoft.github.io/msmarco/ https://microsoft.github.io/msmarco/