4 ms·
Why should TLB flush performance ever be a problem on big machines? You can have one process per core with 128 or more cores, never flush any TLB if you pin tho
by c00lio 3y ago
Why should TLB flush performance ever be a problem on big machines? You can have one process per core with 128 or more cores, never flush any TLB if you pin those processes. And as it is a database, shoveling data from/to disk/SSD is your main concern anyways.
- scottlamb 3y agoPostgreSQL uses synchronous IO, so you won't saturate the CPU with one process (or thread) per core. That said, I think there have been efforts to use io_uring on Linux. I'm not sure how that would work with the process per connection model. Haven't been following it...
- c00lio 3y agoProblem with all kinds of asynchronous I/O is that your processes then need internal multiplexing, akin to what certain lightweight userspace thread models are doing. In the end, it might be harder to introduce than just using OS threads.
- anarazel 3y ago> That said, I think there have been efforts to use io_uring on Linux. I'm not sure how that would work with the process per connection model. Haven't been following it... There's some minor details that are easier with threads in that context, but on the whole it doesn't make much of a difference.
- scottlamb 3y agoI don't understand how it works with thread per connection either. io_uring is designed for systems that have a thread and ring per core, for you to give it a bunch of IO to do at once (batches and chains), and your threads to do other work in the meantime. The syscall cost is amortized or even (through IORING_SETUP_SQPOLL) eliminated. If your code is instead designed to be synchronous and thus can only do one IO at a time and needs a syscall to block on it, I don't think there's much if any benefit in using io_uring. Possibly they'd have a ring per connection and just get an advantage when there's parallel IO going on for a single query? or these per-connection processes wouldn't directly do IO but send it via IPC to some IO-handling thread/process? Not sure either of those models are actually an improvement over the status quo, but who knows.
- anarazel 3y ago> io_uring is designed for systems that have a thread and ring per core That's not needed to benefit from io_uring > for you to give it a bunch of IO to do at once (batches and chains), and your threads to do other work in the meantime. You can see substantial gains even if you just submit multiple IOs at once, and then block waiting for any of them to complete. The cost of blocking on IO is amortized to some degree over multiple IOs. Of course it's even better to not block at all... > If your code is instead designed to be synchronous and thus can only do one IO at a time and needs a syscall to block on it, I don't think there's much if any benefit in using io_uring. We/I have done the work to issue multiple IOs at a time as part of the patchset introducing AIO support (with among others, an io_uring backend). There's definitely more to do, particularly around index scans, but ...
- scottlamb 3y agoOh, I hadn't realized until now I was talking with someone actually doing this work. Thanks for popping into this discussion! > > io_uring is designed for systems that have a thread and ring per core > That's not needed to benefit from io_uring 90% sure I read Axboe saying that's what he designed io_uring for. If it helps in other scenarios, though, great. > Of course it's even better to not block at all... Out of curiosity, is that something you ever want/hope to achieve in PostgreSQL? Many high-performance systems use this model, but switching a synchronous system in plain C to it sounds uncomfortably exciting, both in terms of the transition itself and the additional complexity of maintaining the result. To me it seems like a much riskier change than the process->thread one discussed here that Tom Lane already stated will be a disaster. > We/I have done the work to issue multiple IOs at a time as part of the patchset introducing AIO support (with among others, an io_uring backend). There's definitely more to do, particularly around index scans, but ... Nice. Is the benefit you're getting simply from adding IO parallelism where there was none, or is there also a CPU reduction? Is having a large number of rings (as when supporting a large number of incoming connections) practical? I'm thinking of each ring being a significant reserved block of RAM, but maybe in this scenario that's not really true. A smallish ring for a smallish number of IOs for the query is enough. Speaking of large number of incoming connections, would/could the process->thread change be a step toward having a thread per active query rather than per (potentially idle) connection? To me it seems like it could be: all the idle ones could just be watched over by one thread and queries dispatched. That'd be a nice operational improvement if it meant folks no longer needed a pooler [1] to get decent performance. All else being equal, fewer moving parts is more pleasant... [1] or even if they only needed one layer of pooler instead of two, as I read some people have!