4 ms·
Given the overhead of context switches, is it possible to take a general purpose application like nginx and use a user-mode TCP stack? For instance if I had a n
by inversionOf 11y ago
Given the overhead of context switches, is it possible to take a general purpose application like nginx and use a user-mode TCP stack? For instance if I had a network adapter that is solely dedicated to nginx, and don't need any of the kernel TCP services. Is this even a viable consideration?
I've done high performance nginx, in the million request per second range (there are situations that benefit from these, though unfortunately such discussions always get waylaid by people insisting that performance doesn't matter), but there is enormous system overhead at this rate that I'd like to get around.
- justincormack 11y agoYou can run Nginx with a rump kernel (rumpkernel.org), so with a completely userspace tcp stack. I havent yet done any work on optimising it for 10Gb networking, it is on my TODO list (there are Snabb, Netmap and dpdk drivers, although I might just get it to drive the NIC directly). (Current tests are just with a tap device or raw socket which is very slow).
- jsnell 11y agoIt's possible, you'd need to override the relevant system calls with LD_PRELOADed library. I don't know if a complete drop-in solution is the right solution though. If your application is performance sensitive enough to require embedding a full networking stack, you might as well make use of better APIs. For example it'd be silly to indirect the event dispatching through something poll/select-like. Instead you'd much rather just have the core IO loop call the handlers directly. Or as another example, zero-copy will be impossible with a recv()-like interface where the client provides the buffer that data needs to go to, but will be trivial with an API where it's the network stack giving the client a buffer that already has the data. If you want to experiment with this, mTCP (http://shader.kaist.edu/mtcp/ http://shader.kaist.edu/mtcp/) is probably the right starting point.
- blibble 11y agothis is exactly how openonload works, it's pretty impressive that they get nearly all weird behaviour of the Linux socket API correct (correct behaviour across fork, select/poll/epoll, multicast behaviour, etc). presentation: http://www.openonload.org/openonload-google-talk.pdf http://www.openonload.org/openonload-google-talk.pdf
- acconsta 11y agoThe Arrakis team did it with Haproxy, Reddis, and Memcached, but in a research operating system: http://people.inf.ethz.ch/troscoe/pubs/peter-arrakis-osdi14.pdf http://people.inf.ethz.ch/troscoe/pubs/peter-arrakis-osdi14.... BSD has had userland networking for a while: http://www.bsdcan.org/2014/schedule/events/447.en.html http://www.bsdcan.org/2014/schedule/events/447.en.html I'm not sure if there's anything comparable for Linux.