7 ms·
Linux Kernel vs. DPDK: HTTP Performance Showdown
- tomohawk 4y agoFrom: https://talawah.io/blog/extreme-http-performance-tuning-one-point-two-million/#_2-speculative-execution-mitigations https://talawah.io/blog/extreme-http-performance-tuning-one-... > I am genuinely interested in hearing the opinions of more security experts on this (turning off speculative execution mitigatins). If this is your area of expertise, feel free to leave a comment Are these generally safe if you have a machine that does not have multi-user access and is in a security boundary?
- salmo 4y agoFor regulatory compliance, that’s still not acceptable because it opens the door to cross user data access or privilege escalation. If someone exploited a process with a dedicated unprivileged user, had legit limited access, or got in a container on a physical, they might be able to leverage it for the forces of evil. There’s really no such practical thing as single user Linux. If you’re running a network exposed app without dropping privileges, that’s a much bigger security risk than speculative execution. Now, if you were skipping an OS and going full bare metal, then that could be different. But an audit for that would be a nightmare :).
- flatiron 4y agoThere’s definitely sneakernet Linux boxes. I’ve worked at a bunch of places with random Linux boxes running weird crap totally off the actual network because nobody was particularly sure how to get those programs updated. Technical debt is a pita!
- throwawaylinux 4y agoWhich regulations?
- salmo 4y agoThe worst is the fake regulation that is PCI. But I tried in a SOX audit to avoid buying new gear. I hate auditors. Making sense doesn’t matter. Their interpretation of the control does. I do have a playbook of “how I meet X but not the way you want me to”, but lost that one. Probably spent more $ arguing than the HW cost.
- hedora 4y agoI suspect you are conflating security regulations for Unix users with regulations targeting users of the system. Why would a regulatory framework care if a Linux box running one process was vulnerable to attacks that involve switching UIDs? Converse, why would that same regulatory framework not care if users of that network service were able to impersonate each other / access each others’ data?
- salmo 4y agoMost of the controls are about auditability and data access. But the control frameworks are silly sometimes. Then add in that they’re enforced by 3rd party auditor consultants looking for any reason to drag it out. And yeah, I tried to get this past them for a old singleton system to avoid having to buy a bigger non-standard server.
- urthor 4y agoIt never happens. Anytime the discussion around turning off mitigations comes up, it's trumped by the "why don't we just leave them on and buy more computers" trump card. An easier solution than mitigating is to just upgrade your AWS instance size.
- dsp 4y ago> It never happens. Many intentionally run without mitigations.
- toast0 4y agoYou kind of have to consider the whole system. Who has access, including hypothetical attackers that get in through network facing vulnerabilities, and what privilege escalations they could do with the mitigations on or off. I've ran production systems with mitigations off. All of the intended (login) users had authorized privilege escalation, so I wasn't worried about them. There was a single primary service daemon per machine, and if you broke into that, you wouldn't get anything really useful by breaking into root from there. And the systems where I made a particularly intentional decision were more or less syscall limited; enabling mitigations significantly reduced capacity, so mitigations were disabled. (This was inline with guidance from the dedicated security team where I was working).
- staticassertion 4y agoAs far as I am aware, the capability required to exploit Meltdown is compute and timers. For example, an attacker running a binary directly on the host, or a VM/interpreter (js, wasm, python, etc) executing the attacker's code that exposes a timer in some way. If you just have a dumb "request/response" service you may not have to worry. A database is an interesting case. Scylla uses CQL, which is a very limited language compared to something like javascript or SQL - there's no way to loop as far as I know, for example. I would probably recommend not exposing your database directly to an attacker anyways, that seems like a niche scenario. If you're just providing, say, a gRPC API that takes some arguments, places them (safely) into a Scylla query, and gives you some result, I don't think any of those mitigations are necessary and you'll probably see a really nice win if you disable them. This is my understanding as a security professional who is not an expert on those attacks, because I frankly don't have the time. I'll defer to someone who has done the work. Separately, (quoting the article) > Let's suppose, on the one hand, that you have a multi-user system that relies solely on Linux user permissions and namespaces to establish security boundaries. You should probably leave the mitigations enabled for that system. Please don't ever rely on Linux user permissions/namespaces for a multi-tenant system. It is not even close to being sufficient, with or without those mitigations. It might be OK in situations where you also have strong auth (like ssh with fido2 mfa) but if your scenario is "run untrusted code" you can't trust the kernel to do isolationl > On the other hand, suppose you are running an API server all by itself on a single purpose EC2 instance. Let's also assume that it doesn't run untrusted code, and that the instance uses Nitro Enclaves to protect extra sensitive information. If the instance is the security boundary and the Nitro Enclave provides defense in depth, then does that put mitigations=off back on the table? Yeah that seems fine. > Most people don't disable Spectre mitigations, so solutions that work with them enabled are important. I am not 100% sure that all of the mitigation overhead comes from syscalls, This has been my experience and, based on how the mitigation works, I think that's going to be the case. The mitigations have been pretty brutal for syscalls - though I think the blame should fall on intel, not the mitigation that has had to be placed on top of their mistake. Presumably io_uring is the solution, although that has its own security issues... like an entirely new syscall interface with its own bugs, lack of auditing, no ability to seccomp io_uring calls, no meaningful LSM hooks, etc. It'll be a while before I'm comfortable exposing io_uring to untrusted code.
- fulafel 4y agoBrowser tabs from different origins are equivalent of the multi-user access here, at least in cases where the speculative execution vulnerability can be exploited from JS. Same for other workloads where untrusted parties have a sufficient degree of control on what code executes. Generally.. depends on what you mean by generally. In the casual sense of the word, speculative execution attacks are not very common, so it can be said that most people are mostly safe from them independent of mitigations. Someone might also use "generally safe" to mean proven security against a whole class of attacks, in which case the answer would be no.
- pclmulqdq 4y agoThis was a fascinating read and the kernel does quite nicely in comparison - 66% of DPDK performance is amazing. That said, the article completely nails the performance advantage: DPDK doesn't do a lot of stuff that the kernel does. That stuff takes time. If I recall correctly, DPDK abstractions themselves cost a bit of NIC performance, so it might be interesting to see a comparison including a raw NIC-specific kernel bypass framework (like the SolarFlare one).
- galangalalgol 4y agoIs there a good comparison of these technologies? I've used dpdk for high rate streaming data and it roughly doubled my throughput over 10GE. I hear people using things like dma over Ethernet, and it sounds like there are several competing technologies. My use case is to get something from phy layer into gpu memory as fast as possible, latency is less important than throughput.
- benou 4y agoWhat you're looking for is RDMA. It was mostly restricted to Infiniband (IB) back in the days, but nowadays you probably want RoCEv2. You can look at iWARP too but I think nowadays RoCE won. In any case, the standard software API for RDMA is ibverbs. All adapters supporting RDMA (be it IB, RoCE or iWARP) will expose it. You can get cloud instances with RDMA on AWS and Azure.
- shaklee3 4y agodpdk has rdma/GPUdirect now as well
- medawsonjr 4y agoOr better yet, Mellanox VMA since it's open source (unlike Solarflare OpenOnload) and the NICs are far less expensive.
- bitcharmer 4y ago
- evgpbfhnr 4y agoAt the point you've gotten syscall overhead is definitely going to be a big thing (even without spectre mitigations enabled) -- I'd be very curious to see how far a similar io_uring benchmark would get you. It supports IOPOLL (polling of the socket) and SQPOLL (kernel side polling of the request queue) so hopefully the fact that application driving it is in another thread wouldn't slow it too much... With multi-shot accept/recv you'd only need to tell it to keep accepting connections on the listener fd, but I'm not sure if you can chain recvs to the child fd automatically from kernel or not yet... We live in interesting times!
- JoshTriplett 4y agoI would love to see an io_uring comparison as well; while it's a substantial amount of work to port an existing framework to io_uring, at the point where you're considering DPDK, io_uring seems relatively small by comparison.
- mgaunard 4y agoI personally started a new framework and I went for io_uring for simplicity, that is already giving most of what I need -- asynchronous I/O with no context switching. DPDK is huge and inflexible, it does a lot of things which I'd rather be in control of myself and I think it's easier to just do my own userspace vfio.
- anonymoushn 4y agoIOPOLL is for disks. I think that with very recent kernels you will get busy polling of the socket with just SQPOLL. See here: https://github.com/axboe/liburing/issues/345#issuecomment-1063311635 https://github.com/axboe/liburing/issues/345#issuecomment-10...
- evgpbfhnr 4y agooh! That's not obvious at all from the man page (io_uring_enter.2) If the io_uring instance was configured for polling, by specifying IORING_SETUP_IOPOLL in the call to io_uring_setup(2), then min_complete has a slightly different meaning. Passing a value of 0 instructs the kernel to return any events which are already complete, without blocking. If min_complete is a non-zero value, the kernel will still return immediately if any completion events are available. If no event completions are available, then the call will poll either until one or more completions become available, or until the process has ex‐ ceeded its scheduler time slice. ... Well, TIL -- thanks! and the NAPI patch you pointed at looks interesting too.
- Matthias247 4y agoHi Marc (talawahtech)! Thanks for the exhaustive article. I took a short look at the benchmark setup (https://github.com/talawahtech/seastar/blob/http-performance/apps/tcp_httpd/tcp_httpd.cc https://github.com/talawahtech/seastar/blob/http-performance...), and wonder if some simplifications there lead to overinflated performance numbers. The server here executes a single read() on the connection - and as soon as it receives any data it sends back headers. A real world HTTP server needs to read data until all header and body data is consumed before responding. Now given the benchmark probably sends tiny requests, the server might get everything in a single buffer. However every time it does not, the server will send back two responses to the server - and at that time the client will already have a response for the follow-up request before actually sending it - which overinflates numbers. Might be interesting to re-test with a proper HTTP implementation (at least read until the last 4 bytes received are \r\n\r\n, and assume the benchmark client will never send a body). Such a bug might also lead to a lot more write() calls than what would be actually necessary to serve the workload, or to stalling due to full send or receive buffers - all of those might also have an impact on performance.
- talawahtech 4y agoYea, it is definitely a fake HTTP server which I acknowledge in the article [1]. However based on the size of the requests, and my observation of the number of packets per second in/out being symmetrical at the network interface level, I didn't have a concern about doubled responses. Skipping the parsing of the HTTP requests definitely gives a performance boost, but for this comparison both sides got the same boost, so I didn't mind being less strict. Seastar's HTTP parser was being finicky, so I chose the easy route and just removed it from the equation. For reference though, in my previous post[2] libreactor was able to hit 1.2M req/s while fully parsing the HTTP requests using picohttpparser[3]. But that is still a very simple and highly optimized implementation. FYI, from what I recall, when I played with disabling HTTP parsing in libreactor, I got a performance boost of about 5%. 1. https://talawah.io/blog/linux-kernel-vs-dpdk-http-performance-showdown/#http-server https://talawah.io/blog/linux-kernel-vs-dpdk-http-performanc... 2. https://talawah.io/blog/extreme-http-performance-tuning-one-point-two-million/ https://talawah.io/blog/extreme-http-performance-tuning-one-... 3. https://github.com/h2o/picohttpparser https://github.com/h2o/picohttpparser
- 0xbadcafebee 4y agoFor those like me going "......what is dpdk" The Data Plane Development Kit (DPDK) is an open source software project managed by the Linux Foundation. It provides a set of data plane libraries and network interface controller polling-mode drivers for offloading TCP packet processing from the operating system kernel to processes running in user space. This offloading achieves higher computing efficiency and higher packet throughput than is possible using the interrupt-driven processing provided in the kernel. https://en.wikipedia.org/wiki/Data_Plane_Development_Kit https://en.wikipedia.org/wiki/Data_Plane_Development_Kit https://www.packetcoders.io/what-is-dpdk/ https://www.packetcoders.io/what-is-dpdk/
- jaimex2 4y agoBasically its what people who want to bypass the kernel network stack because they think its slow. They then spend the next few years writing their own stack till they realise they've just re-written what the kernel does and its slower and full of exploits. Yeah, receiving packets is fast when you aren't doing anything with them.
- AlphaSite 4y agoApps probably don’t but networking related code might benefit a lot from DPDK.
- dclusin 4y agoIt’s typically done when end to end latency is more important than full protocol compliance and hardening against myriad types of attackers. High frequency trading is typically where the applications where you see these sorts of implementations being really compelling. For these types of applications there are proprietary implementations that you can buy from vendors that are more suited to latency sensitive applications. The next level of optimization after kernel bypass is to build or buy FPGAs which implement the wire protocol + transport as an integrated circuit.
- grive 4y agoFPGAs are falling out of favor, ASICs are sufficient to implement most network processing. Funnily, one of the biggest DPDK feature is an API to program smartNICs exactly in that way.
- Thaxll 4y agoDo Google and the like actually use TCP in user space or they just use the Linux kernel? Edit: Looks like they do but not for TCP from what I can find: https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/44824.pdf https://static.googleusercontent.com/media/research.google.c...
- jeffinhat 4y agoSee also- Snap: a Microkernel Approach to Host Networking: https://research.google/pubs/pub48630/ https://research.google/pubs/pub48630/
- limoce 4y agoI am not 100% sure that all of the mitigation overhead comes from syscalls, but it stands to reason that a lot of it arises from security hardening in user-to-kernel and kernel-to-user transitions. Will io_uring be also affected by Spectre mitigations given it has eliminated most kernel/user switches? And did anyone do a head-to-head comparison between io_uring and DPDK?
- anonymoushn 4y agoYou can use io_uring with 0 steady-state context switches if you're willing to use 100% CPU on 2 cores :)
- thekozmo 4y agoGood point. This is more of a tcp stack comparison between the kernel and userspace. Seastar has a sharded (per core) stack, which is very beneficial when the number of threads is high
- anonymoushn 4y agoYou can set up one or many rings per core, but the idea I alluded to elsewhere in this comment section of spending 2 cores to do kernel busy polling and userspace busy polling for a single ring is less useful if your alternative makes good use of all cores.
- gonzo 4y agoMy bet is that the stack in VPP is even faster.
- shaklee3 4y agovpp uses dpdk
- thekozmo 4y agoWhat's amazing is that the seastar tcp stack hasn't been changed over the past 7 years, while the kernel received plenty of improvements (in order to close the gap vs kernel bypass mechanisms). Still, for >> 99% of users, there is no need to bypass the kernel.
- touisteur 4y agoI feel this would be a good place to use a spark-based TCP stack. You're bypassing the kernel, have to run stuff as root or risky CAP_ rights, your stack should be as solid as possible. https://www.adacore.com/papers/layered-formal-verification-of-a-tcp-stack https://www.adacore.com/papers/layered-formal-verification-o... Might also give people here some ideas on how to combine symbolic execution, proof, C and SPARK code and how to gain confidence in each part of a network stack. I think there's even some ongoing work climbing up the stack up to HTTP but not sure of the plan (not involved).
- fefe23 4y agoWhy is this interesting to anyone? Haven't we all moved to https by now? Optimizing raw http seems to me like a huge waste of time by now. I say that as someone who has spent years optimizing raw http performance. None of that matters these days.
- staticassertion 4y agoI wouldn't expect HTTPS to make any difference vs HTTP for long lived connections.
- rohith2506 4y agoThis is particularly interesting in HFT where network latency plays a major role in win ratio
- gjulianm 4y agoFor me the interest comes from seeing the speed boost between regular kernel/optimized kernel/DPDK. What you put behind the RX layer doesn't really matter, but it's good to see numbers and things to do when your RX system isn't giving you enough throughput.
- maxgio92 4y agoThank you, very exhaustive and interesting. A note: the link to bftrace script is broken.