3 ms·
It scales poorly because the underlying queue is not lock-free. It was designed in a much simpler time or slower drives (spinning disks). Even if you make it l
by mtanski 10y ago
It scales poorly because the underlying queue is not lock-free. It was designed in a much simpler time or slower drives (spinning disks).
Even if you make it lock-free you still end up paying the communication cost (in/notification). Any time you move data to a different thread that's going to introduce latency (which matters on modern fast drives). Most of these thread pool implementations also suffer head of line blocking... eg. small and fast (page cache) requests get stuck behind long and slow requests.
I know because, I've been fighting this stuff for all of my career. I've recently (in the last two years) tried to push upstream to Linux an a non-blocking page cache read. Thus you have a short circuit case for skipping the queue. In my tests it made samba file sharing ~ 23% faster.
https://lwn.net/Articles/612483/ https://lwn.net/Articles/612483/
https://lwn.net/Articles/636967/ https://lwn.net/Articles/636967/
https://lwn.net/Articles/670231/ https://lwn.net/Articles/670231/
The preadv2 patches stalled... but Christoph revived them and now they are in the kernel. So what's need now is RWF_NONBLOCK.