8 ms·
This is the inaccurate sentence: "Snabbswitch, DPDK and netmap take over the whole network card, not allowing any traffic on that NIC to reach the kernel." Obvi
by s1m0n 11y ago
This is the inaccurate sentence: "Snabbswitch, DPDK and netmap take over the whole network card, not allowing any traffic on that NIC to reach the kernel." Obviously with netmap traffic to the NIC may reach the kernel...
- majke 11y agoThere are many ways to inject packets back to kernel. Tuntap, raw socket on loopback, "dummy" device, etc. So by this count you can always make packets reach the kernel. There are two problems with doing the "take over the nic" techniques: 1) I don't believe you can actually push, say 2M pps back to the kernel with any of this techniques. There is a reason RSS exists, and even if you can process 10M pps on one CPU, it doesn't mean it's easy to insert them back to kernel. 2) I don't think putting a piece of custom code between CloudFlare kernel and network card is feasible on the architectural level. You really want to stand in the way and have to actively forward all these packets?
- s1m0n 11y agoThe title of the article does not mention CloudFlare; only bypassing. The fact that the CloudFlare architecture pushes a higher bandwidth of packets into the network kernel and bypasses the rest does not make it a good technique or to be recommended. If you are primarily interested in the best performance with a single NIC solution then I believe it is suboptimal. Why? You are asking the CPU to do two different types of work; optimized and unoptimized. Because of cache line pollution then the "unoptimized" work via the network kernel will pollute the other work. I may be wrong but I would bet you'd get better performance by separating your CloudFlare specific workload onto two boxes, each with one NIC. In this scenario then no cache line pollution can occur. Of course, these two boxes might not be easily possible within three existing CloudFlare architecture. But this has nothing to do with the general idea of packets bypassing the kernel. After the bypass you want the CPU to process those packets in the most efficient way...