4 ms·
“Can run” and “is useful interactively” are different benchmarks. At this latency, I can still imagine batch or overnight jobs being interesting; for chat, time
by Alisaqqt 2mo ago
“Can run” and “is useful interactively” are different benchmarks. At this latency, I can still imagine batch or overnight jobs being interesting; for chat, time to first useful answer matters much more than whether the weights technically fit. A workload/latency/energy table would make projects like this easier to evaluate.