3 ms·
Yes, 2006 would be a more proper tag. NUMA is interesting, I've had a lot of fun with it on a system to control the adaptive mirrors in ESO's ELT telescope, an
by phkamp 5y ago
Yes, 2006 would be a more proper tag.
NUMA is interesting, I've had a lot of fun with it on a system to control the adaptive mirrors in ESO's ELT telescope, and it is /amazing/ what you can do with a modern server, even on a vanilla kernel.
However, it is a lot harder to exploit NUMA in an application like Varnish, because the requests arrive as a randomized stream, and the effort to figure out which NUMA pool is best situated to handle the request, is not significantly different from just handling the request to begin with.
We used to serve a cache hit in seven system calls, an accept(2), a read(2), a write(2) and four timestamps, and as a result, Varnish servers ran with CPU's >70% idle.
If we had added NUMA-awareness on top of that, we would on average Nnuma-1/Nnuma of the time add at two system-calls to migrate the request to another NUMA domain, and that made little sense.
Since then NICs have become NUMA aware and since we use scatter/gather I/O, that matters a lot more than which CPU core did the overhead processing.
Where NUMA may become mandatory is 100G and 400G networking, but very few people seem interested in running that much traffic through a single server, for reasons of reliability, but I'm keeping an eye on it.
- drewg123 5y agoCheckout TCP_REUSPORT_LB_NUMA on FreeBSD. It will filter incoming TCP connections to listen sockets owned by threads bound to the same NUMA domain as the NIC that received the packets. This is part of the Netflix FreeBSD NUMA work described here https://papers.freebsd.org/2019/eurobsdcon/gallatin-numa_optimizations_network_stack/ https://papers.freebsd.org/2019/eurobsdcon/gallatin-numa_opt... The idea is that you can keep connections local to a NUMA domain. We (Netflix) have a local patch to nginx to support this (basically just setting the socket ioctl on the listen socket after the nginx worker is bound)
- phkamp 5y agoYeah, so there are a lot of things we could optimize for on different kernels, but we prefer to not do so, until we have a really good reason. As I said above, it is not obvious to me that we would gain much over what we already do, but I am keeping an eye on it.
- eqvinox 5y agoWhy did the timestamps require syscalls / do they still do now? (I know CPU counters were a mess back then, but was there any other factor about that?)
- phkamp 5y agoBack then they did. But it is one of those things were I just chose to trust the OS to DTRT, and automatically reaped the benefits of advances "under the hood".