5 ms·
I Have a general vector retrieval question, if you have time to humor me. Suppose I have 10 features per document, each with an embedding. Is it possible to ret
by gregw134 3y ago
I Have a general vector retrieval question, if you have time to humor me. Suppose I have 10 features per document, each with an embedding. Is it possible to retrieve the document with the highest average embedding score across its features? The only approach I can think of is retrieving the top 1k results across each feature to generate a candidate set, then recomputing full scores for each document.
- ashvardanian 3y agoe z! The simplest way with USearch - concatenate 10 embeddings, define a custom metric with Numba, that takes the average of 10 dot-products. Done :)
- gregw134 3y agoCool. What if I want a weighted average of embeddings? Going further, is it possible to adjust the weights at search time?
- ashvardanian 3y agoYes, and yes. The last one may be a bit trickier through Python bindings today, but I can easily include that in the next release… shouldn’t take more than 50 LOC.
- gregw134 3y ago(replying here because hn limits thread length) > Yes, and yes. The last one may be a bit trickier through Python bindings today, but I can easily include that in the next release… shouldn’t take more than 50 LOC. Appreciate it. That'd be game-changing for me. The ultimate thing I'd like to do is actually use a function of the form score = af1(embedding1) + bf2(embedding2) + ... That way you could make adjustments like ignoring feature1 unless its score passes a threshold. I'll try looking at Numba to see if that's possible.
- ashvardanian 3y agoSure, don’t hesitate to reach out to us on Discord. It will be much easier to chat and exchange code snippets there: https://discord.gg/A6wxt6dS9j https://discord.gg/A6wxt6dS9j