3 ms·
You can most definitely batch local models and do unattended inference on a 24/7 basis to maximize utilization on local hardware too. The limits are usually se
by zozbot234 16d ago
You can most definitely batch local models and do unattended inference on a 24/7 basis to maximize utilization on local hardware too. The limits are usually set by some combination of memory utilization for KV cache (particularly on small dGPUs) and overall thermals/power limits (particularly on iGPUs with unified RAM/VRAM). (If you're not near thermal limits, the main alternative to batching is to use MTP or speculative decoding in order to raise arithmetic intensity and speed with the same memory utilization. But batching requests is generally viewed as preferable.) Newer models, especially from DeepSeek, do a nice job of reducing KV cache memory impact for any given context length and/or amount of parallel sessions, so batching on local hw really ought to be quite feasible.