4 ms·
Doesn't seem especially relevant for LLMs. Especially not for training, but even for inference.
by fwip 22d ago
Doesn't seem especially relevant for LLMs. Especially not for training, but even for inference.
- nomel 22d agoSure, but it seems very reasonable to expect some co-hosting of latency sensitive tasks, especially with some low hanging thin client fruit, like cloud gaming and the like. I've been very happy not burning a thousand watts of $0.45/kWh power locally, for my AC to force out of the house, to run blender.