3 ms·
Did you find much difference in inference latency or throughput between baseline and PCA?
by djoldman 2mo ago
Did you find much difference in inference latency or throughput between baseline and PCA?
- stephantul 2mo agoPCA is applied after the model, so there should be no difference in embedding throughput. Lookups in the index should be faster, but that speedup also applies equally to MRL. So I guess the answer is: no