4 ms·
I implemented an highly-scalable user-space TCP stack as part of my Master's thesis [1], last year. One doesn't use an user-space network stack because the Lin
by raphaelj 10y ago
I implemented an highly-scalable user-space TCP stack as part of my Master's thesis [1], last year.
One doesn't use an user-space network stack because the Linux's network stack is slow (it's fast), but because it doesn't scale correctly on an high number of CPUs (> 8 cores) [2]. This is because the kernel suffers from some lock contention when accessing the table containing the socket descriptors.
An user-space stack can be significantly faster when the application layer is really simple (e.g. doing some very simple filtering or routing) and when it does not share any mutable state.
As soon as your application layer starts sharing mutable states between connections (e.g. like a database), you'll start having contention issues similar to those experienced by the kernel, and you'll not gain anything from using an user-space stack. Very often, applications that can benefit from such as stack can also be scaled easily on multiple machines, and it's usually easier to keep using the Linux's stack and add more servers.
--
[1] https://github.com/RaphaelJ/rusty https://github.com/RaphaelJ/rusty
[2] https://github.com/RaphaelJ/rusty/blob/master/doc/img/performances.png https://github.com/RaphaelJ/rusty/blob/master/doc/img/perfor...
- majke 10y ago> This is because the kernel suffers from some lock contention when accessing the table containing the socket descriptors. This would indicate the problem is with packet delivery to application. From my experience even packet delivery to "filter" iptables chain is "slow". But let's assume you are right, can you elaborate? Do you think SO_REUSEPORT on TCP sockets can solve the contention of accept()? https://lwn.net/Articles/542629/ https://lwn.net/Articles/542629/ There are some initiatives improve SO_REUSEPORT CPU affinity, hopefully making it even faster. Update: I misread. Ok, so "table containing the sockets", but this is just a large hash table, nothing too fancy... aRFS for greater locality?
- raphaelj 10y agoA single socket can be shared by multiple cores. That means that the kernel must both protect the socket descriptor from concurrent writes, and can't enforce a TCP link to be handled by a defined core (CPU affinity).
- majke 10y ago> A single socket can be shared by multiple cores Absolutely, it can. But everybody sane avoids that, pinning worker processes to specific CPU's and not sharing sockets between them. The rule of thumb is that spinlocks on the hot path of socket access become a bunch of no-ops if there is no lock contention.
- joosters 10y agoI can't see why pinning is such an obvious choice. The kernel's scheduler may decide that CPU 1 should be woken to handle some new packets, because CPU 2 is busy. If you've pinned the socket to CPU 2, you may be losing out. I get that there are trade-offs between the two modes: pinning can provide better cache usage, you can avoid some locks (but indirectly make the kernel do the work for you) and so on, but I don't see how you can confidently state that pinning is the 'sane' choice. In an ideal world, the kernel has a better overview of the network state and CPU state, and therefore is best positioned to decide which CPU should handle each packet.
- felixgallo 10y agoThe kernel scheduler has no idea which application thread handles which socket. You can formulate an application level plan and then enforce your will with socket and cpu pinning.
- scott_s 10y agoCorrect, but the difficulty is if your application must share the machine with any other application - even short lived ones. That, I think, is what joosters was alluding to. If the assumption that your application is the only consumer of system resources is broken, then you may see pathological scheduling behavior.
- achamayou 10y agoYou can set isolcpus to earmark some cores for your application, and let the kernel manage the rest.
- jhallenworld 10y agoCavium allows you to run multiple instances of the Linux kernel- one pinned to each core of their np (you get to have a window of shared memory between them). It would be interesting to try this on Intel.
- rconti 10y agoTangential, but what are some good resources for understanding the limitations of Linux on larger systems? (8 socket, multi-TB, multi-10gig, NUMA, etc)? Over the years I've found that the trivial questions have saturated the internet and made it very hard to GoogleShoot complicated problems.