4 ms·
For local inference there often isn’t a batch. If I chat with my own llama instance the batch size is one. The model processes a single token at a time doing a
by moconnor 3y ago
For local inference there often isn’t a batch. If I chat with my own llama instance the batch size is one. The model processes a single token at a time doing a lot of vector-matrix multiplication, which is bandwidth bound. CPUs like the M1/2 are very competitive here.
Also, for local inference you only need to be fast enough for many applications. No need to do real time object detection at 1000 FPS or chat at 300 tokens/s (code gen changes this).
- gsuuon 3y agoFor straightforward chat batching wouldn't be very useful, but it can still be very useful for building apps on top of local LLM's which I'm hoping we'll see more and more of.
- liuliu 3y agoSpeculative decoding have higher batch size.
- fomine3 3y agoI understand that many people on HN prefer open-ish LLM on local hardware, but I think it doesn't make sense sadly for efficient hardware usage perspective. Transferring input/output text is almost free cost and local hardware can't be fully utilized by a few people. SaaS make sense here, though I understand that privacy and censorship are matter.