3 ms·
They’re using the FS for caching the KV caches of past requests. It’s why they’re able to charge so little on prompt cache hit.
by cyanf 2y ago
They’re using the FS for caching the KV caches of past requests.
It’s why they’re able to charge so little on prompt cache hit.
- jpgvm 2y agoAhh I missed that. Yes prefix caching and RAG are 2 cases were you will want something like this during inference time.