5 ms·
> This is important because local llm rarely has parallel streams to batch together. I think most people using agent-like usage could easily run any number of
by nullc 4mo ago
> This is important because local llm rarely has parallel streams to batch together.
I think most people using agent-like usage could easily run any number of parallel streams pretty often, but you run out of vram for multiple KV caches, unfortunately.