4 ms·
But why not optimize the TCP part of the kernel?
by Radle 7y ago
But why not optimize the TCP part of the kernel?
- newaccoutnas 7y agoPossibly (I am not a kernel dev) to avoid the memory copy. There are user space network implementations like netmap and dpdk which do this. Possibly this could be optimised in the kernel, I'm not au fait with it enough to be able to say.
- Zekio 7y agoto avoid context switching?
- eptcyka 7y agoTo do that properly, you'll have to interface with the network device directly through userspace. Whilst this is certainly possible, you end up with implementing a device driver in the end which might or might not be what you want to do.
- stingraycharles 7y agoWhich is exactly what Mellanox and others are doing. At QuasarDB, we’re using their drivers to bypass the kernel entirely, which (at scale) saves tremendous amounts of CPU load.
- sanjayts 7y agoAre you talking about these[1]? Also, would you have a rough ballpark figure on the CPU load reduction you have seen using these? [1] http://www.mellanox.com/page/products_dyn?product_family=209&mtag=pmd_for_dpdk http://www.mellanox.com/page/products_dyn?product_family=209...
- stingraycharles 7y agoActually it’s VMA: http://www.mellanox.com/page/software_vma?mtag=vma http://www.mellanox.com/page/software_vma?mtag=vma 100GBit would usually put enormous strains on the CPU (to the point it’s pretty much only doing that), which is taken away entirely by using this. As it turns out, hardware appears to be much faster at these types of things than general purpose CPUs. :)
- raphaelj 7y agoThese user-space TCP stacks usually use an user-space driver or a driver that minimizes system calls with packet aggregation. For example DPDK.
- raphaelj 7y agoI worked on such a TCP stack while doing my MSc. thesis a couple of years ago [1]. Handling TCP in the kernel has some overhead due to system calls. Also, the way sockets are designed does not make them very scalable, as you have lock contention on the TCP state machine. The SO_REUSEPORT feature introduced in Linux 3.9 solve some of these lock issues, but the kernel TCP stack is still not fully parallel [2]. -- [1] https://github.com/RaphaelJ/rusty https://github.com/RaphaelJ/rusty [2] https://raw.githubusercontent.com/RaphaelJ/rusty/master/doc/img/performances.png https://raw.githubusercontent.com/RaphaelJ/rusty/master/doc/...
- derefr 7y ago> But the kernel TCP stack is still not fully parallel Yes, and this project, if moved into kernel-space, would entirely replace that stack and its state-machine. > system calls You’ve still got the overhead (context switches and memory copies) of getting the IP packet out of/into the kernel, which I don’t think is all that much less than the overhead of getting a TCP packet out of/into the kernel. Really, what you want is SR-IOV to allow the user-space process to do direct Ethernet DMA to its own dedicated network card. No copies at all! But if you’re willing to do that, then the application is basically acting as its own kernel... so why not just admit that, and instead of writing a user-space process that has half the features of a kernel, just either 1. write your logic as a Linux kernel driver, or 2. compile your program into a unikernel framework? Then your VM-nee-application’s host can be a proper VMM like Xen or ESXi, where it’s easier to configure that SR-IOV dedication as part of your VM-nee-application’s workload configuration. For this reason, I’ve never understood people trying to do things like this “in user-space.” You’re playing at being a kernel—with all of the problems of being a kernel—without the ability to rely on an existing, well-written kernel as a basis for your logic (like e.g. the parts that handle the L1-L3 layers of the network stack, which you aren’t changing much.)
- scott_s 7y ago> But if you’re willing to do that, then the application is basically acting as its own kernel... so why not just admit that, and instead of writing a user-space process that has half the features of a kernel, just either 1. write your logic as a Linux kernel driver, or 2. compile your program into a unikernel framework? Because the kernel is still a lot of other things for you other than networking - I think it's a stretch to say that all user-space networking makes your work "half" of that of a kernel. And, you're not necessarily the one doing it. You may be an application, and your user-level TCP (including kernel bypass) may be a library from someone else. But to your general point of now you are now well past the city walls, and may run into trouble, I agree. I assume that this sort of thing is only done by a small number of people.
- stingraycharles 7y agoYou’re being downvoted, probably because many people reading this thread already know that kernel <> userland context switching is very expensive, but I think it’s a valid question: there was once a time I also asked this question, and I was happy to learn the answer.
- deleted 7y ago[deleted]