3 ms·
Are you running inference in parallel? 70 tps seems low for parallel execution.
by lostmsu 2mo ago
Are you running inference in parallel? 70 tps seems low for parallel execution.
- jermaustin1 2mo agoIt is on a single 3090, and that seems to be where it averages out. I'll get 85tps on turn 0, but then it settles down to low 70s within a few turns, but holds steady at that. My issue currently is KV Cache, because I can't keep enough parallel caches running (4 is where I'm at), so TTFT (is that the initialism?) can be long when I have a particularly large scene (basically more than 2 NPCs). But my harness does let me offload to any OpenAI compatible endpoint, I just prefer local cuz free.