19 ms·
Linux network performance parameters
- deleted 3y ago[deleted]
- gjulianm 3y agoThis is great, not just the parameters themselves but all the steps that a packet follows from the point it enters the NIC until it gets to userspace. Just one thing to add regarding network performance: if you're working in a system with multiple CPUs (which is usually the case in big servers), check NUMA allocation. Sometimes the network card will be in one CPU while the application is executing on a different one, and that can affect performance too.
- suprjami 3y agoPackagecloud have a great article series which goes into much more detail and with code study. If you really want to learn the network send and receive, these are the articles to read: https://blog.packagecloud.io/monitoring-tuning-linux-networking-stack-receiving-data/ https://blog.packagecloud.io/monitoring-tuning-linux-network... https://blog.packagecloud.io/monitoring-tuning-linux-networking-stack-sending-data/ https://blog.packagecloud.io/monitoring-tuning-linux-network... https://packagecloud.io/blog/illustrated-guide-monitoring-tuning-linux-networking-stack-receiving-data/ https://packagecloud.io/blog/illustrated-guide-monitoring-tu...
- mikece 3y ago[flagged]
- Thaxll 3y agoThis is kind of an urban legend, do you think the multi millions servers from Google, Amazon etc... have those performance issues?
- jeffbee 3y agoThe big guys don't have the patience to wait for Linux kernel networking to be fast and scalable. They bypass the kernel and take over the hardware. https://blog.acolyer.org/2019/11/11/snap-networking/ https://blog.acolyer.org/2019/11/11/snap-networking/
- corbet 3y agoThat's funny ... the "big guys" are some of the biggest contributors to the Linux network stack, almost as if they were actually using it and cared about how well it works.
- jeffbee 3y agoHistory has shown that tons of Linux networking scalability and performance contributions have been rejected by the gatekeepers/maintainers. The upstream kernel remains unsuitable for datacenter use, and all the major operators bypass or patch it.
- sophacles 3y agoAll the major operators sometimes bypass or patch it for some use cases. For others they use it as is. For other still they laugh at you for taking the type of drugs that makes one think any CPU is sufficient to handle networking in code. Networking isn't a one size fits all thing - different networks have different needs, and different systems in any network will have different needs. Userland networking is great until you start needing to deal with weird flows or unexpected traffic - then you end up either needing something a bit more robust and your performance starts dropping because you added a bunch of branches to your code or switched over to a kernel implementation that handles those cases. I've seen a few cases of userland networking being slower than just doing the kernel - and being kept because sometimes the what you care about is control over packet lifecycle more than raw throughput. Kernels prioritize robust network stacks that can handle a lot of cases good enough. Different implementations handle different scenarios better - there's plenty of very high performance networking done with vanilla linux and vanilla freebsd.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- sophacles 3y agoPerformance parity on which axis? For which use case? Talking generally about "network performance" is approximately as useful as talking generally about "engine performance". Just like it makes no sense to compare a weed-eater engine to a locomotive diesel without talking about use case and desired outcomes, it makes no sense to compare "performance of FreeBSD network stack" and "Linux network stack" without understanding the role those systems will be playing in the network. Depending on context, FreeBSD, Linux or various userland stacks can be a great, average, or terrible choices.
- circularfoyers 3y agoCan you provide some examples of different contexts where Linux or FreeBSD might be better or worse choices?
- sophacles 3y agoSure: Linux is a networking swiss army knife (or maybe a dremmel). It can do a lot of stuff reasonably well. It has all sorts of knobs and levers, so you can often configure it to do really weird stuff. I tend to reach for it first to understand the shape of a problem/solution. BSD is fantastic for a lot of server applications, particularly single tenant high throughput ones like mail servers, dedicated app servers, etc. A great series of case studies have come out of Netflix on this (google for "800Gbps on freebsd netflix" for example - every iteration of that presentation is fantastic and usually discussed here at least once, and Drew G. shows up in comments and answers questions). It's also pretty nice for firewalling/routing small and medium networks - (opn|pf)sense are both great systems for this built on FreeBSD (apologies for the drama injection this may cause below). One of the reasons I reach for linux first unless I already know the scope and shape of the problem is that the entire "userland vs kernel" distinction is much blurrier there. Linux allows you to pass some or all traffic to userland at various points in the stack and in various ways, and inject code at the kernel level via ebpf, leading to a lot of hybrid solutions - this is nice in middleboxes where you want some dynamism and control, particularly in multi-tenant networks (and thats the space my work is in, so it's what I know best) Please bear in mind that these are my opinions and uses/takes on the tools. Just like with programming there's a certain amount of "art" (or maybe "craft") to this, and other folks will have different (but likely just a valid) views - there's a lot of ways to do anything in networking.
- nolist_policy 3y agoFrom every benchmark I've seen so far, Linux has always been faster than the BSDs. For example, look at these benchmarks from 2003[1]. Makes you wonder where the myth comes from. The newest benchmark I could find[2] points in the same direction. Does anyone have more recent data? [1] http://bulk.fefe.de/scalability/ http://bulk.fefe.de/scalability/ [2] https://matteocroce.medium.com/linux-and-freebsd-networking-cbadcdb15ddd https://matteocroce.medium.com/linux-and-freebsd-networking-...
- dekhn 3y agoLinux has better network performance than FreeBSD over nearly every use case I've seen.
- lacrosse_tannin 3y agoLet me know when the BSDs have a complete networking stack. for example: - SO_BINDTODEVICE - an option for not-symmetric NAT
- raggi 3y agoThis largely depends on how the application is written and in particular what if any non-POSIX interfaces it uses. If you are looking to hit line rates with UDP, or looking to head well above ~1-10gbps with TCP, you're fast headed into territory where you likely need to move away from POSIX. (For a super dumb benchmark, oversized buffers amortizing syscall overhead might get you to 10gbps on TCP, but in a real application everything changes) Once you're headed over 10gbps you'll quickly run into a hard need to retune things even for TCP, earlier if you're talking to non-local hosts. Once you're over 25gbps you're headed into the territory where you'll need to fix drivers, fix cpu tuning and so on. For a recent real world example: when we were doing performance analysis of our offloading patches for Tailscale we identified problems with the current default CPU frequency scaler for Intel CPUs on current kernels, and reached out to the maintainers with data.
- tptacek 3y agoIn a thread that is about tuning Linux network performance, a comment like this is essentially arson.
- deleted 3y ago[deleted]
- freedomben 3y agoCould anyone recommend a video or video series covering similar material? There lots on networking in general, but I've had a hard time finding some on Linux specific implementation
- deleted 3y ago[deleted]
- 8K832d7tNmiQ 3y agoI'm also seconding this, but from microcontroller perspective. I want to try developing a simple tcp echo server for a microcontroller, but most examples just use the vendor's own tcp library and put no effort explaining how to manually setup and establish connection to the router.
- patmorgan23 3y agoWell you can always read the standard
- hanikesn 3y agoBecause implementing TCP not just as a toy is incredibly difficult with tripwires that even subject matter experts struggle to get it right without decades of in field testing.
- doctorpangloss 3y agoDoes performance tuning for Wi-Fi adapters matter? On desktops, other than disabling features, can anything fix the problems with i210 and i225 ethernet? Those seem to be the two most common NICs nowadays. I don't really understand why common networking hardware and drivers are so flawed. There is a lot of attention paid to RISC-V. How about start with a fully open and correct NIC? They'll shove it in there if it's cheaper than an i210. Or maybe that's impossible.
- jeffbee 3y agoi225 is just broken but I get excellent performance from i210. 1gb is hardly challenging on a contemporaneous CPU, and the i210 offers 4 queues. What's your beef with i210?
- trustingtrust 3y agoThere are 3 revisions of i225 and Intel essentially got rid of it and launched i226. That one also seems to be problematic [1] . Why is it exponentially harder to make a 2.5gbps NIC when the 1gbps NIC (i210 and i211) has worked well for them. Shouldn't it be trivial to make it 2.5x? They seem to make good 10gbps NICs so I would assume 2.5gbps shouldn't need a 5th try from intel ? [1] - https://shorturl.at/esCNP https://shorturl.at/esCNP
- jeffbee 3y agoThe bugs I am aware of are on the PCIe side. i225 will lock up the bus if it attempts to do PTM to support PTP. That's a pretty serious bug. You would think Intel has this nailed since they invented PCIe and PCI for that matter. Apparently not. Maybe they outsourced it.
- uep 3y agoThis is really interesting for me to read. I encountered a DMA lockup in the hardware by an Ethernet MAC implementation on an ARM chip. It was a Synopsys Designware MAC implementation. It would specifically lockup when PTP was enabled. From my testing, it seemed like it would specifically lockup if some internal queue was overrun. This was speculation on my part, because it would only lockup if I tried to enable timestamping on all packets. It seemed to work alright if the hardware filter was used to only timestamp PTP packets. This can be a significant limitation though, as it can prevent PTP from working with VLANs or DSA switch tags, since the hardware can't identify PTP packets with those extra prefixes. The PTP timestamps would arrive as a separate DMA transaction after the packet DMA transaction. It very possibly could have been poor integration into the ARM SOC, but your PTP-specific issue on x86 makes me wonder.
- klabb3 3y agoA random thing I ran into with the defaults (Ubuntu Linux): - net.ipv4.tcp_rmem ~ 6MB - net.core.rmem_max ~ 1MB So.. the tcp_rmem value overrides by default, meaning that the TCP receive window for a vanilla TCP socket actually goes up to 6MB if needed (in reality - 3MB because of the halving, but let's ignore that for now since it's a constant). But if I "setsockopt SO_RCVBUF" in a user-space application, I'm actually capped at a maximum 1MB, even though I already have 6MB. If I try to reduce it from 6MB to e.g. 4MB, it will result in 1MB. This seems very strange. (Perhaps I'm holding it wrong?) (Same applies to SO_SNDBUF/wmem...) To me, it seems like Linux is confused about the precedence order of these options. Why not have core.rmem_max be larger and the authoritative directive? Is there some historical reason for this?
- pengaru 3y agonet.ipv4.tcp_rmem max is a limit for the auto-tuning the kernel performs once you do SO_RCVBUF the auto-tuning is out of the picture for that socket, and net.core.rmem_max becomes the max. It's pretty clearly documented @ Documentation/networking/ip-sysctl.rst Edit: downvotes, really? smh
- dekhn 3y agoAnd to add: the kernel autotunes better than you can, so leave that enabled unless you're Vint Cert, Jim Gettys, or Vern Paxton.
- mjan22640 3y agoChanged my name, thanks for the tip!
- LukeShu 3y ago1. While your context about auto-tuning is accurate and valuable, it doesn't really address the fundamental strangeness that the parent post is commenting about: It's still strange that it can auto-tune to a higher value than you can manually tune it to. 2. It's always valuable to provide further references, but I'd guess that down-voters found the "It's pretty clearly documented" phrasing a little condescending? Perhaps "See the docs at [] for more information."? 3. "Please don't comment about the voting on comments. It never does any good, and it makes boring reading."
- napkin 3y agoJust changing Linux's default congestion control (net.ipv4.tcp_congestion_control) to 'bbr' can make a _huge_ difference in some scenarios, I guess over distances with sporadic packet loss and jitter, and encapsulation. Over the last year, I was troubleshooting issues with the following connection flow: client host <-- HTTP --> reverse proxy host <-- HTTP over Wireguard --> service host On average, I could not get better than 20% theoretical max throughput. Also, connections tended to slow to a crawl over time. I had hacky solutions like forcing connections to close frequently. Finally switching congestion control to 'bbr' gives close to theoretical max throughput and reliable connections. I don't really understand enough about TCP to understand why it works. The change needed to be made on both sides of Wireguard.
- drewg123 3y agoThe difference is that BBR does not use loss as a signal of congestion. Most TCP stacks will cut their send windows in half (or otherwise greatly reduce them) at the first sign of loss. So if you're on a lossy VPN, or sending a huge burst at 1Gb/s on a 10Mb/s VPN uplink, TCP will normally see loss, and back way off. BBR tries to find Bottleneck Bandwidth rate. Eg, the bandwidth of the narrowest or most congested link. It does this by measuring the round trip time, and increasing the transmit rate until the RTT increases. When the RTT increases, the assumption is that a queue is building at the narrowest portion of the path and the increase of RTT is proportional to the queue depth. It then drops rate until the RTT normalizes due to the queue draining. It sends at that rate for a period of time, and then slightly increases the rate to see if RTT increases again (if not, it means that the queuing that saw before was due to competing traffic which has cleared). I upgraded from a 10Mb/s cable uplink to 1Gb/s symmetrical fiber a few years ago. When I did so, I was ticked that my upload speed on my corp. VPN remained at 5Mb/s or so. When I switched to RACK TCP (or BBR) on FreeBSD, my upload went up by a factor of 8 or so, to about 40Mb/s, which is the limit of the VPN.
- heybrendan 3y agoYou seem quite knowledgeable in this domain. Have you authored any blog posts to expand on this topic? I would welcome the chance to learn more from you.
- cryptonector 3y agoNothing about PMTUD?
- zamadatix 3y agoFor TCP sockets I'd rather just MSS clamp on the internet gateway. On top of too many things just dropping PMTUD, enabling it results in a slower process while MSS clamping hijacks the initial TCP open messages directly.
- cryptonector 3y agoThere's PMTUD that doesn't depend on ICMP.
- zamadatix 3y agoFor TCP in Linux the only thing I know of is net.ipv4.tcp_mtu_probing=2 which is still slower than clamping at the edge. You can also run into weird slowdowns in cases with packet loss even after the initial discovery. If you don't have a way to clamp but absolutely need the interface to have jumbo enabled for local traffic performance it's probably the best fallback but even then I'm not sure it's worth the extra headache it causes.
- toast0 3y agoFWIW, I put together a PMTUD test site you might find interesting http://pmtud.enslaves.us/ http://pmtud.enslaves.us/
- cryptonector 3y agoNice URL!
- raggi 3y agoThis doc kinda needs to say "TCP" somewhere, as it's very focused on TCP concerns - which is useful, people are mostly using TCP. The default UDP tunings are awfully low and as such are notably missing.
- leshow 3y agoDo you have any good resources for UDP tuning?
- BitPirate 3y agoThe typical painpoint is the low maximum buffer sizes. net.core.rmem_max net.core.wmem_max e.g. wireguard-go will hit those limits of not executed with CAP_NET_ADMIN.
- sophacles 3y agoI have gotten quite a bit of mileage out of this slide deck: https://events.static.linuxfound.org/sites/events/files/slides/LinuxConJapan2016_makita_160712.pdf https://events.static.linuxfound.org/sites/events/files/slid... It's older so some details have changed over time, but the concepts are still relevant. It also has a lot of useful search terms to get you started.
- zartstrom 3y agoI enjoyed skimming through the article. Very well researched and presented. But can anybody help me out, who tunes linux network parameters on a regular basis?
- teleforce 3y agoGreat overview of the Linux network queues as provided in the Figure, should paste it on the wall somewhere. Brendan's System Performance books provide nice coverage on Linux network performance and more [1]. It's already in the second edition, both are excellent books but the 2nd edition focuses mainly on Linux whereas the 1st edition also include Solaris. There's also a more recent book on BPF Performance Tools by him [2]. [1] Systems Performance: Enterprise and the Cloud, 2nd Edition (2020) https://www.brendangregg.com/systems-performance-2nd-edition-book.html https://www.brendangregg.com/systems-performance-2nd-edition... [2] BPF Performance Tools: https://www.brendangregg.com/bpf-performance-tools-book.html https://www.brendangregg.com/bpf-performance-tools-book.html