5 ms·
Excellent article ! I got hit by the exact same issue which is described in the fermilab paper, namely packet reordering caused by intel drivers. It took me se
by Agingcoder 5y ago
Excellent article !
I got hit by the exact same issue which is described in the fermilab paper, namely packet reordering caused by intel drivers. It took me several days to diagnose the problem. Interestingly enough, the problem virtually disappeared when running tcpdump, which, after a lot of reading on the innards of the linux TCP stack, and prodding with ebpf, eventually led me to conjecture that it was a scheduling/core placement issue. Pinning my process clearly made the problem disappear, and then finding the paper nailed it.
Networks are not my specialty (I come from a math background, am self taught, and had always dismissed them as mere plumbing) , but I have to say that I came out of this difficult (for me) investigation with a great appreciation for networking in general, and now enjoy reading anything I can find about them.
It's never too late to learn, and I have yet to find something in software engineering which is not interesting once you take a closer look at it!
- SaveTheRbtz 5y ago> tcpdump ... linux TCP stack ... ebpf > Networks are not my specialty I wish all network non-specialist were like you!
- Natfan 5y agoI have nowhere near any of these skills, and I know I'm not a specialist.
- brohee 5y agoI think the problem disappearing while running tcpdump is one of the truest instances of Schrödingbug...
- Agingcoder 5y agoIt was perfectly reproducible. The original symptom was very low throughput, which is what prompted the investigation . Without tcpdump, low throughput and high reordering, with tcpdump, high throughput ( which is why I couldn't figure out what was going on). I'd be very interested if someone with kernel experience could tell me what's specific about tcpdump.
- o-__-o 5y agoPcap captures are not multithreaded so you are pinning to a single core. This entire thread is interesting because it is highlighting a similar problem with my virtualized router. When I pinned the router vm to specific cpus the problem goes away. I switched to openstack which doesn’t give me the best control over cpu capabilities and the problem has manifested in a worse form. My uninformed opinion is that there are underlying concurrency problems with multithreaded user land-kernel interaction and some nic drivers (consumer intel and Broadcom hardware)
- Agingcoder 5y agoAh ! Thanks for this. Being single threaded does not prevent you from having your process being migrated from one core to another though, no? Or do you mean that pcap captures are pinned?
- baruch 5y agoNetworking actually has tons of interesting and complex math. Congestion control is a rabbit hole of math and control theory.
- Agingcoder 5y agoIndeed. I ended up reading quite a bit about congestion control while investigating a different issue (sending data from a 25gb box to a 1gb one over a 10gb/100ms latency link didn't work well since the bigger nic would saturate the smaller one which then dropped packets, and this caused the tcp window to shrink significantly ), and it was extremely interesting. The whole problem of having multiple agents competing with different strategies and incomplete information to maximize their network throughput also reminded me of economics. Tcp pacing essentially solved the problem (and if not available brutally traffic shape).