5 ms·
BMX: A Freshly Baked Take on BM25
- leobg 2y ago> Entropy-weighted similarity: We adjust the similarity scores between query tokens and related documents based on the entropy of each token. Sounds a lot like BM25 weighted word embeddings (e.g. fastText).
- dmezzetti 2y ago> Sounds a lot like BM25 weighted word embeddings (e.g. fastText). If you're interested in this topic, I wrote an article on this method back in 2020: https://medium.com/towards-data-science/building-a-sentence-embedding-index-with-fasttext-and-bm25-f07e7148d240 https://medium.com/towards-data-science/building-a-sentence-...
- nblgbg 2y agoThanks for sharing the article. How exactly are you combining BM25 and fastText? Are you combining the TF-IDF score + embedding distance? What are the weights for each of these?
- flawn 2y agoGemischtes Brot!
- bernihackernews 2y agobaguetter library for the win!
- deepsquirrelnet 2y agoVery cool! Glad to see continued research in this direction. I’ve really enjoyed reading the Mixedbread blog. If you’re interested in retrieval topics, they’re doing some cool stuff.
- timsuchanek 2y agoAmazing! When will we have this in the major databases?
- yokee 2y agoSuper cool! It is definitely a good choice for the RAG system.
- antman 2y agoHow about computational complexity? There seems to be a small improvement in metrics but not sure if it is enough to switch to bmx
- intalentive 2y agoNeat. I wonder how GPT-4’s query expansion might compare with SPLADE or similar masked BERT methods. Also if you really want to go nuts you can apply term expansion to the document corpus.
- herrmannfield 2y agowhy do i have this vibe ? https://blogs.perficient.com/2012/09/25/a-mathematical-model-for-assessing-page-quality/ https://blogs.perficient.com/2012/09/25/a-mathematical-model...