4 ms·
I may have this wrong, but here's my best understanding of it: Ruby supports multi-threading, but unless you're using the new (and experimental) Ractor feature
by drewbug01 4y ago
I may have this wrong, but here's my best understanding of it:
Ruby supports multi-threading, but unless you're using the new (and experimental) Ractor feature, you're subject to the global interpreter lock in most cases (with a few important and useful exceptions, like some kinds of I/O). That means that Ruby servers will typically employ multi-processing in addition to (or in place of) multi-threading as a way to increase performance and use multiple CPUs - otherwise, multiple threads just end up competing for the global interpreter lock and the additional threads don't increase performance as much as you would hope, especially if serving those requests requires any actual work to be done in Ruby code.
Puma supports a multi-processing mode, where a main Puma process forks multiple workers (each running multiple threads), and each worker listens on the same socket. The linux kernel distributes the load between the workers, and then the workers distribute the load internally between their threads. Since the global interpreter lock is a per-process thing, this is a pretty effective way to get more throughput for a Ruby server.
The problem is that you can't directly control how the kernel is going to balance incoming requests across the multiple workers listening on a socket. Because Ruby does support some instances where threads can run concurrently - like network I/O - it's possible that the kernel may end up handing off multiple requests to one worker process when there were others that were idle and could have handled the request. Doesn't sound like a big deal - but because most threaded Ruby operations do not run concurrently that means that the actual Ruby code that needs to process that request is going to be competing for the global interpreter lock.
So basically this allows a worker process that is already handling requests to insert a tiny delay before accepting another one - which gives an idle worker process a chance to accept it instead. On balance, this means that you'll get higher utilization of the CPU resources available to you and will often result in a lower average latency for all requests.
The PR that added this in Puma 5.0 is here: https://github.com/puma/puma/pull/2079 https://github.com/puma/puma/pull/2079
- sytse 4y agoThanks so much for the clear write up and link! Cool to see it is my fellow GotLab team member Kamil who added this.