4 ms·
The KV cache is a great example. When you own hardware, vllm etc can keep your KV cache warm indefinitely for bursty large token loads vs clouds will purge quit
by drewnick 2mo ago
The KV cache is a great example. When you own hardware, vllm etc can keep your KV cache warm indefinitely for bursty large token loads vs clouds will purge quite often.
One of my tasks has a very large but very static system prompt and instructions, think something like 128k tokens. On our own hardware, we keep that cached, sprinkle in the 8k of goodness needed for output, and DSV4 Flash 0731 can generate our output at ~50 tps on $10k worth of hardware.
I've been benchmarking vs publics clouds (which are tbf insanely cheap!) and it's basically a break even in 24-36 months if nothing changes, which it could for better or worse.
Combined with the security and stability of internal hardware, to me it's in "no brainer" territory for this workload, even if models don't improve.
Our only "risk" is better cheaper hardware or cloud costs, which given the trajectory is murky at best short of a bust.