5 ms·
I worked on such a TCP stack while doing my MSc. thesis a couple of years ago [1]. Handling TCP in the kernel has some overhead due to system calls. Also, the
by raphaelj 7y ago
I worked on such a TCP stack while doing my MSc. thesis a couple of years ago [1].
Handling TCP in the kernel has some overhead due to system calls.
Also, the way sockets are designed does not make them very scalable, as you have lock contention on the TCP state machine. The SO_REUSEPORT feature introduced in Linux 3.9 solve some of these lock issues, but the kernel TCP stack is still not fully parallel [2].
--
[1] https://github.com/RaphaelJ/rusty https://github.com/RaphaelJ/rusty
[2] https://raw.githubusercontent.com/RaphaelJ/rusty/master/doc/img/performances.png https://raw.githubusercontent.com/RaphaelJ/rusty/master/doc/...
- derefr 7y ago> But the kernel TCP stack is still not fully parallel Yes, and this project, if moved into kernel-space, would entirely replace that stack and its state-machine. > system calls You’ve still got the overhead (context switches and memory copies) of getting the IP packet out of/into the kernel, which I don’t think is all that much less than the overhead of getting a TCP packet out of/into the kernel. Really, what you want is SR-IOV to allow the user-space process to do direct Ethernet DMA to its own dedicated network card. No copies at all! But if you’re willing to do that, then the application is basically acting as its own kernel... so why not just admit that, and instead of writing a user-space process that has half the features of a kernel, just either 1. write your logic as a Linux kernel driver, or 2. compile your program into a unikernel framework? Then your VM-nee-application’s host can be a proper VMM like Xen or ESXi, where it’s easier to configure that SR-IOV dedication as part of your VM-nee-application’s workload configuration. For this reason, I’ve never understood people trying to do things like this “in user-space.” You’re playing at being a kernel—with all of the problems of being a kernel—without the ability to rely on an existing, well-written kernel as a basis for your logic (like e.g. the parts that handle the L1-L3 layers of the network stack, which you aren’t changing much.)
- scott_s 7y ago> But if you’re willing to do that, then the application is basically acting as its own kernel... so why not just admit that, and instead of writing a user-space process that has half the features of a kernel, just either 1. write your logic as a Linux kernel driver, or 2. compile your program into a unikernel framework? Because the kernel is still a lot of other things for you other than networking - I think it's a stretch to say that all user-space networking makes your work "half" of that of a kernel. And, you're not necessarily the one doing it. You may be an application, and your user-level TCP (including kernel bypass) may be a library from someone else. But to your general point of now you are now well past the city walls, and may run into trouble, I agree. I assume that this sort of thing is only done by a small number of people.
- yxhuvud 7y agoIt does raise the question about if it would be possible by the kernel to implement features that so to say erect city walls in a way that processes can't access whatever it want. Ie, granting access to some of the ethernet DMA but not all of it, or something similar. Perhaps it isn't possible in theory, but perhaps it is, or could be.
- toast0 7y agoIf you align the interface's packet hashing with kernel and application space cpu pining, most of the tcp locking is limited to a single cpu, so there's no cross cpu contention. Microsoft calls this receive side scaling, and it's also available in FreeBSD. On FreeBSD, this really helps with tcp data packets, but connection setup still has bottlenecks; there's not an api for setting up outgoing sockets to align with the cpu you're on, but it's possible to do it with manual port assignment on outgoing connections.
- chx 7y ago> the kernel TCP stack is still not fully parallel Every mainstream OS today predates mainstream SMP. While support is built into each, it's just ... support. They just were not made in the world where every system is SMP. To see how big a design difference that makes, well, look at Erlang or Go.
- nickpsecurity 7y agoOr BeOS or DragonflyBSD.
- waddlesplash 7y agoTechnically DragonFlyBSD was the evolution of a non-SMP kernel, though indeed they've reorganized enough that it more than qualifies as "engineered for SMP." I think the Haiku kernel was also? At any rate it certainly is now, in much the way DragonFlyBSD's was if nothing else.
- nickpsecurity 7y agoYou're starting to make me think most of what I read about BeOS isn't true. One of those things was that it was designed for strong, concurrency support. They bragged about it in the demo showing how it didn't degrade under massive load. Another paper mentioned the "benaphores" that improved on semaphores a bit. So, was it not designed for concurrency and they just somehow fixed a lot of it later? That retrofits usually don't work is why I believed claims that concurrency was a design goal.
- waddlesplash 7y agoConcurrency and SMP are two very different things. One can have a kernel that handles concurrency better than anything else alive but it does not have any support for SMP. But ... I didn't say anything about BeOS in my comment? I was talking about Haiku, which has very different origins than BeOS, especially on the kernel front where the internal architecture was and is pretty different from the Be kernel. I actually don't know how much the Be kernel was designed for SMP from the first days; I think it was but I'm not sure. At any rate it definitely did have better concurrency for desktop usage than anything else at the time, I believe.