3 ms·
Running inference requires sharing intermediate matrix results between nodes. Faster networking speeds that up.
by elcritch 7mo ago
Running inference requires sharing intermediate matrix results between nodes. Faster networking speeds that up.
- wokkel 7mo agoI read (but cannot find this anymore) that the information sent from layer to layer is minimal. The actual matrix work happens within a layer. They are not doing matrix multiplication over the netwerk (that would be insane latency wise).
- elcritch 7mo agoThe LLM/transformers attention layers require an O(n^2) operation between all tokens, which does require significant bandwidth. Yes the latency hurts performance, that why it’s only achieving ~8tok/s.