2 ms·
Nice! I’ve been working on something similar and found similar results. In my experiments, I used lots of embedding models and the results were not nearly as u
by stephantul 2mo ago
Nice! I’ve been working on something similar and found similar results.
In my experiments, I used lots of embedding models and the results were not nearly as uniform as this curve, just FYI. I didn’t use any of the API-based models though
I also wrote about this exact comparison when using PCA and MRL to quantize static models, see: https://stephantul.github.io/blog/mrl-pca/ https://stephantul.github.io/blog/mrl-pca/
- dcastm 2mo agoThank you! Will take a look at your results. I couldn't find much when I first looked into this, which is why I ended up writing the article.
- stephantul 2mo agoAh I meant more to say that I was working on this as well. I haven’t published the results for this comparison specifically yet.
- djoldman 2mo agoDid you find much difference in inference latency or throughput between baseline and PCA?
- stephantul 2mo agoPCA is applied after the model, so there should be no difference in embedding throughput. Lookups in the index should be faster, but that speedup also applies equally to MRL. So I guess the answer is: no