2 ms·
> Though only 5gig Ethernet? Can’t they do usb-c / thunderbolt 40 Gb/s connections like Macs? Does the network speed matter that much when TFA talks about outp
by TacticalCoder 7mo ago
> Though only 5gig Ethernet? Can’t they do usb-c / thunderbolt 40 Gb/s connections like Macs?
Does the network speed matter that much when TFA talks about outputting a few tens of tokens per second? Ain't 5 Gbit/s plenty for that? (I understand the need to load the model but that'd be local already right?)
- elcritch 7mo agoRunning inference requires sharing intermediate matrix results between nodes. Faster networking speeds that up.
- wokkel 7mo agoI read (but cannot find this anymore) that the information sent from layer to layer is minimal. The actual matrix work happens within a layer. They are not doing matrix multiplication over the netwerk (that would be insane latency wise).
- elcritch 7mo agoThe LLM/transformers attention layers require an O(n^2) operation between all tokens, which does require significant bandwidth. Yes the latency hurts performance, that why it’s only achieving ~8tok/s.