4 ms·
Why not bump it to 10,000 threads? The post shows: the OS scheduler struggles badly, 18x slower allocation, 17x slower context switching. That’s measured overhe
by earcar 1y ago
Why not bump it to 10,000 threads? The post shows: the OS scheduler struggles badly, 18x slower allocation, 17x slower context switching. That’s measured overhead, not theory.
Complexity? We migrated in 30 minutes. It’s just Async blocks, not goroutine scheduling gymnastics.
Not claiming it’s a silver bullet - the post explicitly says “use threads for CPU work”. But for I/O-bound LLM streaming, the massive improvement is real and in production.