3 ms·Accelerate CPU Based LLM Inference with a Vector Index on the Output Embeddings1 points by dithered_djinn 2y ago