3 ms·
This article a bit fishy to me. E.g.: > In almost all network stacks, per-flow IPsec processing is serialized due to implementation specific constraints of upp
by __michaelg 11y ago
This article a bit fishy to me. E.g.:
> In almost all network stacks, per-flow IPsec processing is serialized due to implementation specific constraints of upper level protocols that require in-order packet processing and lower level hardware interfaces that are optimized for latency. The lack of concurrency, combined with memory copy requirements, inline cryptographic operations and multiple traversals of the network layer are to blame for poor performance.
How is that different for the TLS solution? You still need somewhat in-order processing of the TLS stream, memory copies won't really go away with an additional TLS proxy, and the network layer traversals also aren't really that different. Sure one can put work into optimizing all of that, but that appears orthogonal to using TLS to me.
What follows it the usual collection IPSec management downsides -- which may be true but shouldn't affect performance.
IMO the interesting part is pretty far down the post:
> Additionally, TLS record sizes are often larger than MTU (16KB vs 1500/9000B), resulting in lower setup costs for encapsulation.
This may be the actual important change that affects the throughput. It would be interesting to see the effect on real workloads, i.e. not in synthetic benchmarks. Vice versa, it would also be interesting to see the same benchmarks with TLS record sizes limited to the usual MTU sizes.
As the saying goes: I'm not impressed if your prototype is slightly better than what we have in production for years.
- bbjnicklin 11y agoMichaelg, Regarding the network, the paths are different in a meaningful way, you might want to trace the paths for yourself to see how: https://upload.wikimedia.org/wikipedia/commons/3/37/Netfilter-packet-flow.svg https://upload.wikimedia.org/wikipedia/commons/3/37/Netfilte.... Specifically, with IPSec, you make another pass through the network layer, after hitting XFRM in the protocol layer. This is not the case for a normal flow. This “loop” is significant because of the frequency with which it happens. Also, as the post states, random memory accesses are the killer. This additional loop and the logic within exacerbate the problem. Also, If you’ve been using IPsec in production for multiple years, it’s possible that you using non-AESNI optimized ciphers, so the speedup could in fact be MUCH greater. If you can share a bit about your specific deployment, we would be happy to provide guidance.