3 ms·
Slop language aside, the abstract insight is directionally correct I think. CPUs were already becoming much more important during training for test-time scalin
by spmurrayzzz 2mo ago
Slop language aside, the abstract insight is directionally correct I think.
CPUs were already becoming much more important during training for test-time scaling, but there you were still bottlenecked by GPU compute since the gradient updates back to the policy model are the actual gating factor.
During normal inference though, CPUs are becoming more of a bottleneck for more advanced workloads. Even if you have 20 agents running in parallel, if they're all compiling Rust concurrently your total wall-clock time per task is no longer bound by the decode throughput of the upstream model. You're just waiting for tools to execute. This gets compounded by VM/container overhead as well if you're doing the totally local sandbox approach.
- mohamedkoubaa 2mo agoIt's been said before that the best human programmers were always bottlenecked by compile test loop time