5 ms·
It doesn't mention in the article why the kernel is so slow at processing network packets. I'm not a kernel programmer so this may be utterly wrong, but wouldn'
by ilanco 11y ago
It doesn't mention in the article why the kernel is so slow at processing network packets. I'm not a kernel programmer so this may be utterly wrong, but wouldn't it be possible to sacrifice some feature for speed by disabling it in the kernel code?
- zobzu 11y agotldr: cloudflare like to reimplement things and claim it solves the world problems. Thats cool, thats how open source work. Fork, copy, reimplement, try new stuff. that-said...slightly-longer: for what they do there are alternatives like ipset for other things its not as clear-cut, hence things like PF_RING. its not that great thought, you're sacrificing all features for fast sniffing. technically a good zero-copy implementation of packet mmap w/ a userspace ring would achieve +- the same thing, too.
- lukego 11y agoSmart people are working on this: https://lwn.net/Articles/629155/ https://lwn.net/Articles/629155/ However, even after a few years now I don't think they are seeing light at the end of the tunnel.
- zurn 11y agoNote they cite higher numbers as the current status than the Cloudfare article. In the LWN article it says "The kernel, today, can only forward something between 1M and 2M packets per core every second" Also eg. http://vger.kernel.org/netconf2009_slides/LinuxCon2009_JesperDangaardBrouer_final.pdf http://vger.kernel.org/netconf2009_slides/LinuxCon2009_Jespe... from 6 years ago show forwarding at 4 Mpps. (Forwarding is, of course, receiving AND sending instead of just receiving so these should translate to higher RX-only numbers.)
- acdha 11y agoOne thing to remember is that high-speed packet filtering is an unusual workflow and CloudFlare operates at a much greater scale than most of us see: most Linux devices are not connected to 10G, much less 100G, networks and they're usually doing more work than looking at a packet to decide whether to accept or reject it. The fact that the APIs and the kernel stack were designed many years before those kind of speeds were possible doesn't matter because most sites don't have that much traffic and most server applications will bottleneck at doing other work well before that point. The example in the article found a single core handling 1.4M packets per second. If you're running a web-server shoveling data out to clients those packets are going to be close to the maximum size which, if I haven't screwed up the math, looks something like this: 1.4M * 1400 bytes (assuming a low MTU) * 8 (bytes -> bits) = 15Gbps That's not to say that there isn't still plenty of room to improve and, as lukego noted, there's a lot of work in progress (see e.g. https://lwn.net/Articles/615238/ https://lwn.net/Articles/615238/ on work to batch operations to avoid paying some of the processing costs for every packet) but for the average server you'd find bottlenecks on something like a database, application logic, request handling, client network capacity, etc. before the network stack overhead is your greatest challenge. The people who encounter this tend to be CDN vendors like CloudFlare and security people who need to filter, analyze, or generate traffic on levels which are at least the the scale of a large company (e.g. https://github.com/robertdavidgraham/masscan https://github.com/robertdavidgraham/masscan).
- lukego 11y agoPeople are also spending hundreds of billions of dollars on equipment like routers ever year. These could be Linux boxes if the kernel had sensible performance. I am kind of amazed that the Linux kernel did not become the dominant data-plane for the networking industry ahead of proprietary implementations from Cisco, Juniper, etc. Hopefully Snabb Switch will have better luck there... ;-)
- asdfaoeu 11y agoASICs are always going to be way faster than a Linux kernel.
- __d 11y agoMost switch/router dataplane processing is done in hardware, with control by custom drivers in the control-plane OS. Cisco IOS is variously hosted on a proprietary real-time OS, QNX or Linux. Juniper JunOS is FreeBSD-based. Arista AOS is Linux-based. With the advent of SDN, the packet-rate limitations of a general-purpose OS lead to things like DPDK and Snabb, but they both run within a Linux host environment (DPDK can use FreeBSD as well; not sure re: Snabb).
- scurvy 11y agoSorry if I sound mean, but this is just a long apologist post about how things are just so hard. Really? Why? Why can't Linux match BSD's performance? Also, 10gb servers are not rare by any means. Take a look around the next time you walk in a colo. 10 gb servers everywhere.
- justincormack 11y agoWell the BSDs have issues too. FreeBSD/Netflix has been doing in kernel SSL to try to fix some issues, and there is Netmap available for packet processing in userspace. So they are running into similar issues as Linux.
- gonzo 11y ago
- zurn 11y agoIt isn't as slow as they claim, just like was discussed in HN comments to the predecessor article ("How to receive a million packets per second"). Still they repeat the general claim of "Vanilla Linux can do only about 1M pps". Makes for better headlines I guess.
- jeffreyrogers 11y agoDoesn't look like anyone has really answered your question yet. There are three main reasons the kernel is slow for networking: per-packet dynamic memory allocation, lots of memory copying, and system call overheads. The first two can be improved by modifying the kernel, and I think people are attempting to do this. The system call overheads arise naturally from having the networking code in the kernel. Basically every time you perform a system call the kernel has to save the userspace context, do the system call, and restore the context. This takes time and is bad for cache locality. But as others noted, for most people who aren't cloudflare this doesn't really matter.
- acconsta 11y ago>But as others noted, for most people who aren't cloudflare this doesn't really matter. Aren't most web applications I/O bound? The Arrakis team sped up Memcached and Haproxy quite a lot by bypassing the kernel. It seems like there could be a large market for these techniques as they become easier to use. http://people.inf.ethz.ch/troscoe/pubs/peter-arrakis-osdi14.pdf http://people.inf.ethz.ch/troscoe/pubs/peter-arrakis-osdi14....
- jeffreyrogers 11y agoHmmm, that's a good point. I was thinking more of a typical webapp, but you're probably right that there are certain classes of applications (e.g. caching) that are I/O bound under load.
- s1m0n 11y agoReason #4: Another problem is that the network kernel was never designed to do internet on the mass scale desired today. Companies like whatsapp devoted lots of time to getting e.g. 2M concurrent TCP connections (considered good) running on a single box, mainly because of the greedy overhead and design of the legacy network kernel. Whereas, in theory it should be possible to have 10M or more concurrent TCP connections on modern average hardware. So from this POV then the legacy network kernel is the bloated memory greedy mess that Java is to software development. See http://c10m.robertgraham.com/p/manifesto.html http://c10m.robertgraham.com/p/manifesto.html