14 ms·
Extreme HTTP Performance Tuning
- 120bits 5y agoVery well written. - I have nodejs server for the APIs and its running on m5.xlarge instance. I haven't done much research on what instance type should I go for. I looked up and it seems like c5n.xlarge(mentioned in the article) is meant compute optimized. That cost difference isn't much between m5.xlarge and c5n.xlarge. So, I'm assuming that switching to c5 instance would be better, right? - Does having ngnix handle the request is better option here? And setup reverse proxy for NodeJS? I'm thinking of taking small steps on scaling an existing framework.
- talawahtech 5y agoThanks! The c5 instance type is about 10-15% faster than the m5, but the m5 has twice as much memory. So if memory is not a concern then switching to c5 is both a little cheaper and a little faster. You shouldn't need the c5n, the regular c5 should be fine for most use cases, and it is cheaper. Nginx in front of nodejs sounds like a solid starting point, but I can't claim to have a ton of experience with that combo.
- deleted 5y ago[deleted]
- nodesocket 5y agom5 has more memory, if you application is memory bound stick with that instance type. I'd recommend just using a standard AWS application load balancer in front of your Node.js app. Terminate SSL at the ALB as well using certificate manager (free). Will run you around $18 a month more.
- danielheath 5y agoFor high level languages like node, the graviton2 instances offer vastly cheaper cpu time (as in, 40%). That’s the m6g / c6g series. As in all things, check the results on your own workload!
- jeffbee 5y agoVery nice round-up of techniques. I'd throw out a few that might or might not be worth trying: 1) I always disable C-states deeper than C1E. Waking from C6 takes upwards of 100 microseconds, way too much for a latency-sensitive service, and it doesn't save you any money when you are running on EC2; 2) Try receive flow steering for a possible boost above and beyond what you get from RSS. Would also be interesting to discuss the impacts of turning off the xmit queue discipline. fq is designed to reduce frame drops at the switch level. Transmitting as fast as possible can cause frame drops which will totally erase all your other tuning work.
- xtacy 5y agoI suspect that the web server's CPU usage will be pretty high (almost 100%), so C-state tuning may not matter as much? EDIT: also, RSS happens on the NIC. RFS happens in the kernel, so it might not be as effective. For a uniform request workload like the one in the article, statically binding flows to a NIC queue should be sufficient. :)
- duskwuff 5y agoDoes C-state tuning even do anything on EC2? My intuition says it probably doesn't pass through to the underlying hardware -- once the VM exits, it's up to the host OS what power state the CPU goes into.
- jeffbee 5y agoIt definitely works and you can measure the effect. There's official documentation on what it does and how to tune it: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/processor_state_control.html https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/processo...
- duskwuff 5y agoOkay, so it looks as though it only applies to certain large instance types -- presumably ones which are large enough that it makes sense for the host to statically allocate CPU cores (or even sockets) to a guest. Interesting.
- alufers 5y agoThat is one hell of a comprehensive article. I wonder how much impact would such extreme optimizations on a real-world application, which for example does DB queries. This experiment feels similar to people who buy old cars and remove everything from the inside except the engine, which they tune up so that the car runs faster :).
- 101008 5y agoYes, my experience (not much) is that what makes YouTube or Google or any of those products really impressive is the speed. YouTube or Google Search suggestion is good, and I think it could be replicable with that amount of data. What is insane is the speed. I can't think how they do it. I am doing something similar for the company I work on and it takes seconds (and the amount of data isn't that much), so I can't wrap my head around it. The point is that doing only speed is not _that_ complicated, and doing some algorithms alone is not _that_ complicated. What is really hard is to do both.
- ecnahc515 5y agoA lot of this is just spending more money and resources to make it possible to optimize for speed. With sufficient caching with and a lot of parallelism makes this possible. That costs money though. Caching means storing data twice. Parallelism means more servers (since you'll probably be aiming to saturate the network bandwidth for each host). Pre-aggregating data is another part of the strategy, as that avoids using CPU cycles in the fast-path, but it means storing even more copies of the data! My personal anecdotal experience with this is with SQL on object storage. Query engines that use object storage can still perform well with the above techniques, even though querying large amounts of data from object is slow. You can bypass the slowness of object storage if you pre-cache the data somewhere else that's closer/faster for recent data. You can have materialized views/tables for rollups of data over longer periods of time, which reduces the data needed to be fetched and cached. It also requires less CPU due to working with a smaller amount of pre-calculated data. Apply this to every layer, every system, etc, and you can get good performance even with tons of data. It's why doing machine-learning in real- is way harder than pre-computing models. Streaming platforms make this all much easier as you can constantly be pre-computing as much as you can, and pre-filling caches, etc. Of course, having engineers work on 1% performance improvements in the OS kernel, or memory allocators, etc will add up and help a lot too.
- drenvuk 5y agoI'm of two minds with regards to this: This is cool but unless you have no authentication, data to fetch remotely or on disk this is really just telling you what the ceiling is for everything you could possibly run. As for this article, there are so many knobs that you tweaked to get this to run faster it's incredibly informative. Thank you for sharing.
- joshka 5y ago> this is really just telling you what the ceiling is That's a useful piece of info to know when performance tuning a real world app with auth / data / etc.
- Thaxll 5y agoThere was a similar article from Dropbox years ago: https://dropbox.tech/infrastructure/optimizing-web-servers-for-high-throughput-and-low-latency https://dropbox.tech/infrastructure/optimizing-web-servers-f... still very relevant
- secondcoming 5y agoFantastic article. Disabling spectre mitigations on all my team's GCE instances is something I'm going to check out. Regarding core pinning, the usual advice is to pin to the CPU socket physically closest to the NIC. Is there any point doing this on cloud instances? Your actual cores could be anywhere. So just isolate one and hope for the best?
- halz 5y agoPinning to the physically closest core is a bit misleading. Take a look at output from something like `lstopo` [https://www.open-mpi.org/projects/hwloc/ https://www.open-mpi.org/projects/hwloc/], where you can filter pids across the NUMA topology and trace which components are routed into which nodes. Pin the network based workloads into the corresponding NUMA node and isolate processes from hitting the IRQ that drives the NIC.
- ricktdotorg 5y agowow, i had wondered about pinning in the cloud. this is a fantastic tip - thank you!
- brobinson 5y agoThere are a bunch more mitigations that can be disabled than he disables in the article. I usually refer to https://make-linux-fast-again.com/ https://make-linux-fast-again.com/
- diroussel 5y agoDid you consider wrk2? https://github.com/giltene/wrk2 https://github.com/giltene/wrk2 Maybe you duplicated some of these fixes?
- talawahtech 5y agoYea, I looked it wrk2 but it was a no-go right out the gate. From what I recall the changes to handle coordinated omission use a timer that has a 1ms resolution. So basically things broke immediately because all requests were under 1ms.
- skyde 5y agoso twrk doesn't handle coordinated omission or you found a different way to do it?
- talawahtech 5y agoI didn't make any coordinated omission changes (I really didn't make many changes in general), so twrk does what wrk does. It attempts to correct it after the fact by looking for requests that took twice as long as average and doing some backfilling[1]. I am no expert where coordinated omission is concerned, but my understanding is that it is most problematic in scenarios where your p90+ latency is high. Looking at the results for the 1.2M req/s test you have the following latencies: p50 203.00us p90 236.00us p99 265.00us p99.99 317.00us pMAX 626.00us If you were to apply wrk's coordinated omission hack to these result, the backfilling only starts for requests that took longer than p50 x 2 (roughly) = 406us, which is probably somewhere between p99.999 and pMAX; a very, very small percentage. I am not claiming that wrk's hack is "correct", just that I don't think coordinated omission is a major concern for *this specific workload/environment* 1. https://github.com/wg/wrk/blob/a211dd5a7050b1f9e8a9870b95513060e72ac4a0/src/stats.c#L33 https://github.com/wg/wrk/blob/a211dd5a7050b1f9e8a9870b95513...
- throwdbaaway 5y agoIf I understand correctly, coordinated omission handling only matters if the benchmark is done with a fixed rate RPS right? In this case, it looks like a closed model benchmark where a fixed number of client threads just go as fast as they can. edit: Oh, perhaps wrk2 still relies on the timer even when not specifying a fixed rate RPS.
- specialist 5y agoWhat is the theoretical max req/s for a 4 vCPU c5n.xlarge instance?
- talawahtech 5y agoThere is no published limit, but based on my tests the network device for the c5n.xlarge has a hard limit of 1.8M pps (which translates directly to req/s for small requests without pipelining). There is also a quota system in place, so even though that is the hard limit, you can only operate at those speeds for a short time before you start getting rate-limited.
- specialist 5y agoImproving from 12.4% to 66.6% of theoretical max is kinda amazing. Presented this way may help noobs like me with capacity planning.
- strawberrysauce 5y agoYour website is super snappy. I see that it has a perfect lighthouse score too. Can you explain the stack you used and how you set it up?
- talawahtech 5y agoIt is a statically generated site created with vitepress[1] and hosted on Cloudflare Pages[2]. The only dynamic functionality is the contact form which sends a JSON request to a Cloudflare Worker[3], which in turn dispatches the message to me via SNS[4]. It is modeled off of the code used to generate Vue blog[5], but I made a ton of little modifications, including some changes directly to vitepress. Keep in mind that vitepress is very much an early work in progress and the blog functionality is just kinda tacked on, the default use case is documentation. It also definitely has bugs and is under heavy development so wouldn't recommend it quite yet unless you are actually interested in getting your handa dirty with Vue 3. I am glad I used it because it gave me an excuse to start learning Vue, but unless you are just using the default theme to create a documentation site, it will require some work. 1. https://vitepress.vuejs.org/ https://vitepress.vuejs.org/ 2. https://pages.cloudflare.com/ https://pages.cloudflare.com/ 3. https://workers.cloudflare.com/ https://workers.cloudflare.com/ 4. https://aws.amazon.com/sns/ https://aws.amazon.com/sns/ 3. https://github.com/vuejs/blog https://github.com/vuejs/blog
- remram 5y agoOn the other hand you could probably make the table of content be always visible when the screen size allows it. Clicking on the burger in the site menu to get a page-specific sidebar is a bit counter-intuitive.
- strawberrysauce 5y agoThanks :). Found one flaw in your already crazy optimized vitpress site - the images aren't cached :P
- ricktdotorg 5y agocf-cache-status: HIT
- 0xbadcafebee 5y agoVery well written, bravo. TOC and reference links makes it even better.
- bigredhdl 5y agoI really like the "Optimizations That Didn't Work" section. This type of information should be shared more often.
- fabioyy 5y agodid you try DPDK?
- miohtama 5y agoHow much head room there would be if one were to use Unikernel and skip the application space altogether?
- the8472 5y agoSince it's CPU-bound and spends a lot of time in the kernel would compiling the kernel for the specific CPU used make sense? Or are the CPU cycles wasted on things the compiler can't optimize?
- talawahtech 5y agoRecompiling the kernel using profile guided optimizations[1] is yet another thing on the (never-ending) to-do list. 1. https://lwn.net/Articles/830300/ https://lwn.net/Articles/830300/
- ta988 5y agoCould you make a profile of just a bunch of functions on a running system?
- ArtWomb 5y agoWow. Such impressive bpftrace skill! Keeping this article under my pillow ;) Wonder where the next optimization path leads? Using huge memory pages. io_uring, which was briefly mentioned. Or kernel bypass, which is supported on c5n instances as of late...
- ta988 5y agoKernel bypass?
- fierro 5y agoHow can you be sure the estimated max server capability is not actually just a limitation in the client, i.e, the client maxes out at sending 224k requests / second. I see that this is clearly not the case here, but in general how can one be sure?
- 0xEFF 5y agoUse N clients. Increase N until you’re sure.
- mh- 5y agoYou parallelize the load from multiple clients (running on separate hardware). There are some open source projects that facilitate this sort of workload (and the subsequent aggregation of results/stats.)
- trashcan 5y agohttps://locust.io/ https://locust.io/ is a good example
- paracyst 5y agoI don't have anything to add to the conversation other than to say that this is fantastic technical writing (and content too). Most of the time, when similar articles like this one are posted to company blogs, they bore me to tears and I can't finish them, but this is very engaging and informative. Cheers
- talawahtech 5y agoThanks, that actually means a lot. It took a lot of work, not just on the server/code, but also the writing. I asked a lot of people to review it (some multiple times) and made a ton of changes/edits over the last couple months. Thanks again to my reviewers!
- 3gg 5y agoVery educational and well-written, thank you.
- londons_explore 5y agoSome of these things could be fixed upstream and everyone see real perf gains... For example, having dhclient (a very popular dhcp client) leave open an AF_PACKET socket causing a 3% slowdown in incoming packet processing for all network packets seems... suboptimal! Surely it can be patched to not cause a systemwide 3% slowdown (or at least to only do it very briefly while actively refreshing the DHCP lease)?
- talawahtech 5y agoI would also love to see that dhclient issue resolved upstream, or at least a cleaner way to work around it. But we should also be mindful that for most workloads the impact is probably way, way less. Some of these things really only show up when you push things to their extremes, so it probably just wasn't on the developer's radar before.
- lttlrck 5y agoI believe systemd-networkd has its own implementation of DHCP and therefore doesn't use dhclient. But I wonder if it's behavior is any better in this respect. This has piqued my interest.
- mercora 5y agosystemd-networkd keeps open that kind of socket for LLDP but apparently not for the DHCP client code. wpa_supplicant also keeps open this type of socket on my local system. and the dhcpd daemons on my routers have some of those too for each interface... i wonder if the slow path here could be avoided by using separate network namespaces in a way these sockets don't even get to see the packets...
- lttlrck 5y agoInteresting. Looks like LLDP can be switched off in the network config. https://systemd.network/systemd.network.html https://systemd.network/systemd.network.html
- zdw 5y agoI wonder what the results would be if all the optimizations were applied except for the security-related mitigations, which were left enabled.
- nhoughto 5y agoI’d love to have the time (and ability!) to do this level of digging. Amazing write up to, very well presented.
- brendangregg 5y agoGreat work, thanks for sharing! Systems performance at its best. Nice to see the use of the custom palette.map (I forget to do that myself and I often end up hacking in highlights in the Perl code.) BTW, those disconnected kernel stacks can probably be reconnected with the user stacks by switching out the libc for one with frame pointers; e.g., the new libc6-prof package.
- talawahtech 5y agoThank you for sharing all your amazing tools and resources brendangregg! I wouldn't have been able to do most of these optimizations without FlameGraph and bpftrace. I actually did the same thing and hacked up the perl code to generate the my custom palette.map Thanks for the tip re: the disconnected kernel stacks. They actually kinda started to grow on me for this experiment, especially since most of the work was on the kernel side.
- anarazel 5y agoIs libc6-prof just glibc recompiled with -fno-omit-frame-pointer? I did that a couple times and found that while that fixes a few system calls, it doesn't fix all of them. I think the main issue was several syscalls being called from asm, which wasn't unsurprisingly isn't affected by -fno-omit-frame-pointer.
- brendangregg 5y agoRight, it is. It fixed my hot-path syscalls on x86 (via read/write functions, pthread_mutex functions, etc.). But if you have syscalls called via asm outside of libc (by who?) then they need frame pointers as well.
- MichaelMoser123 5y agoWhen is it advisable to turn off spectre/meltdown mittigations in practice? My guess is that if you are on a server and not running any user supplied code then you are on the safe side; on condition that you could exclude buffer overuns by running managed code/java or by using Rust.
- toast0 5y agoSo the unspoken part of your question is when is it useful to turn off mitigations. The answer to that is when your application makes a lot of syscalls / when syscalls are a bottleneck beyond the actual work of the syscalls. This case, where it's all connection handling and serving a small static piece of data is a clear example; there's almost no userland work to be done before it goes to another syscall so any additional cost for the user/kernel barrier is going to hurt. Then the question becomes who can run code on your server; also condidering maybe there's a remote code execution vulnerability in your code, or library code you use. Is there a meaningful barrier that spectre/meltdown mitigations would help enforce? Or would getting RCE get control over everything of substance anyway?
- MichaelMoser123 5y agoif you have an event driven system then end up with very frequent system calls.
- anarazel 5y agoPartially that can be amortized with io_uring... At the cost of some complexity, of course.
- MichaelMoser123 5y agoio_uring was added to linux 5.1, that was in 2019. I have to admit that i didn't yet have the chance to use it. https://en.wikipedia.org/wiki/Io_uring https://en.wikipedia.org/wiki/Io_uring Did you use io_uring? Is its performance much better than or comparable with using aio_read/aio_write for block io? (i did use async io for block io).
- jart 5y ago> Disabling [spectre] mitigations gives us a performance boost of around 28% Every couple months these last several years there always seems to be some bug where the fix only costs us 3% performance. Since those tiny performance hits add up over time, security is sort of like inflation in the compute economy. What I want to know is how high can we make that 28% go? The author could likely build a custom kernel that turns off stuff like pie, aslr, retpoline, etc. which would likely yield another 10%. Can anyone think of anything else?
- ronsor 5y agoMost of these mitigations are worse than useless in an environment not executing untrusted code. Simply put, if you have a dedicated server and you aren't running user code, you don't need them.
- vlz 5y agoBut of course other exploits (e.g. in your webapp) might lead to "running user code" where you didn't expect it and then the mitigations could prevent privilege escalation, couldn't they?
- vladvasiliu 5y agoBut if you have a dedicated server for your web app, if there's some kind of exploit in it allowing for random code to be run, said code already has access to everything it needs, right? The interesting data will probably be whatever secrets the app handles, say database credentials, so the attacker is off to the races. They probably don't care about having root in particular.
- ex_amazon_sde 5y ago> if there's some kind of exploit in it allowing for random code to be run, said code already has access to everything it needs On the same host there could be SSL certificates, credentials in a local MTA, credentials used to run backups and so on. Or the application itself could be made of multiple components where the vulnerable one is sandboxed.
- micropoet 5y agoImpressive stuff
- injinj 5y agoGreat work, thanks! I'm curious whether disabling the slow kernel network features competes with an tcp bypass stack. I did my own wrk benchmark [0], but I did not try to optimize the kernel stack beyond pinning CPUs and busypoll, because the bypass was about 6 times as fast. I assumed that there is no way the kernel stack could compete with that. This article shows that I may be wrong. I will definitely check out SO_ATTACH_REUSEPORT_CBPF in the future. [0] https://github.com/raitechnology/raids/#using-wrk-httpd-loading https://github.com/raitechnology/raids/#using-wrk-httpd-load...
- talawahtech 5y agoThat is an area I am curious about as well, especially if you throw io_uring into the mix. I think most kernel bypass solutions get some of their gains by just forcing you to use the same strategies covered in the perfect locality section. It doesn't all just come from the "bypass" part. Even if isn't quite as fast as DPDK and co, it might be close enough for some people to start opting to stick with the tried and true kernel stack instead of the more exotic alternatives.
- injinj 5y agoMy gut feeling with io_uring is that it wouldn't help as much with messaging applications with 100 byte request/reply patterns. It would be better in a with a pipelined situation, through a load balancing front end. I would love to be proven wrong, though.
- talawahtech 5y ago1.2M req/s means 2.4M (send/recv) syscalls per second. I definitely think io_uring will make a difference. Just not sure if it will be 5% or 25%.
- throwdbaaway 5y ago> EC2 X-factor? > Even after taking all the steps above, I still regularly saw a 5-10% variance in performance across two seemingly identical EC2 server instances > To work around this variance, I tried to use the same instance consistently across all benchmark runs. If I had to redo a test, I painstakingly stopped/started my server instance until I got an instance that matched the established performance of previous runs. We notice similar performance variance when running benchmark on GCP and Azure. In the worst case, there can be a 20% variance on GCP. On Azure, the variance between identical instances is not as bad, perhaps about 10%, but there is an extra 5% variance between normal hours and off-peak hours, which further complicates things. It can be very frustrating to stop/start hundreds of times for hours to get back an instance with the same performance characteristic. For now, I use a simple bash for-loop that checks the "CPU MHz" value from lscpu output, and that seems to be reliable enough.
- jiggawatts 5y agoWhy would you expect two different virtual machines to have identical performance? I would expect that just the cache usage characteristics of "neighbouring" workloads alone would account for at least a 10% variance! Not to mention system bus usage, page table entry churn, etc, etc... If you need more than 5% accuracy for a benchmark, you absolutely have to use dedicated hosts. Even then, just the temperature of the room would have an effect if you leave Turbo Boost enabled! Not to mention the "silicon lottery" that all overclockers are familiar with... This feels like those engineering classes where we had to calculate stresses in every truss of a bridge to seven figures, and then multiply by ten for safety.
- throwdbaaway 5y agoI didn't expect identical performance, but a 10~20% variance is just too big. For example, if https://www.cockroachlabs.com/guides/2021-cloud-report/ https://www.cockroachlabs.com/guides/2021-cloud-report/ got a "slow" GCP virtual machine but a "fast" azure virtual machine, the final result could totally flip. The more problematic scenario, as mentioned in the article, is when you need to do some sort of performance tuning that can take weeks/months to complete. On the cloud, you either have to keep the virtual machine running all the time (and hope that a live migration doesn't happen behind the scene to move it to a different physical host), and do the painful stop/start until you get back the "right" virtual machine before proceeding to do the actual work. We discovered this variance a couple of months ago. And this article from talawah.io is actually the first time I have seen anyone else mentioning about it. It still remains a mystery, because we too can't figure out what contributes to the variance using tools like stress-ng, but the variance is real when looking at MySQL commits/s metric. > If you need more than 5% accuracy for a benchmark, you absolutely have to use dedicated hosts. After this ordeal, I am arriving at that conclusion as well. Just the perfect excuse to build a couple of ryzen boxes.
- habibur 5y agoThat can be done with HTTP. But right now it's all HTTPS specially when you are serving APIs over the Internet. And once I switch to HTTPS I see a dramatic drop in throughput like x10. A http 15k req/sec drops down to 400 req/sec once I start serving it over HTTPS. I see no solution to it as everything has to https now.
- astrange 5y agoHTTPS especially TLS1.3 is not slow. x86 has had AES acceleration since 2010. It might need different tuning or you might be negotiating a slow cipher.
- ComputerGuru 5y agoThe SSL handshake (which affects TTFB) isn’t AES.
- astrange 5y agoRight, but TLS1.3 improves that especially with 0RTT. Before that you had things like session resumption for repeat clients, or if your server was overloaded you could use an external HTTPS proxy.
- jcelerier 5y ago> HTTPS especially TLS1.3 is not slow. are you saying that the parent poster is dreaming when he sees his performance divided by 37 when turning on https ?
- mtoddsmith 5y agoAt a previous job they tracked down some slow https performance in a game server to OpenSSL lib allocating/reallocation new buffers for each zip’d request. Patching that gave a huge performance increase and saved them from buying some fancy $500k hardware to offload the https processing.
- hinkley 5y agoI can still remember the days when /dev/random slowed down SSL session handshakes.
- baybal2 5y agoTake a note, no quick cheat like DPDK was used. This shows you can make a regular Linux program using Linux network stack to approach something handcoded with DPDK.
- SaveTheRbtz 5y agoThe analysis itself is quite impressive: a very systematic top-down approach. We need more people doing stuff like this! But! Be careful applying tunables from the article "as-is"[1]: some of them would destroy TCP performance: net.ipv4.tcp_sack=0 net.ipv4.tcp_dsack=0 net.ipv4.tcp_timestamps=0 net.ipv4.tcp_moderate_rcvbuf=0 net.ipv4.tcp_congestion_control=reno net.core.default_qdisc=noqueue Not to mention that `gro off` that will bump CPU usage by ~10-20% on most real world workload, Security Team would be really against turning off mitigations, and usage of `-march=native` will cause a lot of core dumps in heterogenous production environments. [1] This is usually the case with single purpose micro-benchmarks: most of the tunables have side effects that may not be captured by a single workflow. Always verify how the "tunings" you found on the internet behave in your environment.
- bbeausej 5y agoThank you for the amazing article and detailed insights. Great writing style and approaches. How long did you spend researching this subject to produce such an in depth report?
- talawahtech 5y agoHard to say exactly. I have been working on this in my spare time, but pretty consistently since covid-19 started. A lot of this was new to me, so it wasn't all as straight-forward as it seems in the blog. As a ballpark I would say I invested hundreds of hours in this experiment. Lots of sidetracks and dead ends along the way, but also an amazing learning experience.
- cakoose 5y agoThis was great! Reminds me a lot of this classic CS paper: Improving IPC by Kernel Design, by Jochen Liedke (1993) https://www.cse.unsw.edu.au/~cs9242/19/papers/Liedtke_93.pdf https://www.cse.unsw.edu.au/~cs9242/19/papers/Liedtke_93.pdf
- alinspired 5y agoWhat was an MTU in the test, how increasing it affects the results ? Reminds me how complicated it was to generate 40Gbit/sec of http traffic (with default MTU) to test F5 Bigip appliances, luckily TCL irules had `HTTP::retry`
- talawahtech 5y agoThe MTU is 9001 within the VPC, but the packets are less than 250 bytes so the MTU doesn't really come into play. This test is more about packets/s than bytes/s.
- Bellamy 5y agoI have done some performance optimization but this article has 30% stuff I have never heard of. Great work and thanks!
- truth_seeker 5y agoVery impressive analysis. Thanks for sharing.
- linlin1991 5y agofalse null
- HugoDaniel 5y ago"Many of these specific optimizations won't really benefit you unless you are already serving more than 50k req/s to begin with."
- volta83 5y agoI'm missing one thing from the article, that is commonly missing from performance-related articles. When you talk about playing whack-a-mole with the optimizations, this is what you are missing: > What's the best the hardware can do? You don't say in the article. The article only says that you start at 250k req/s, and ends at 1.2 req/s. Is that good? Is your optimization work done? Can you open a beer and celebrate? The article doesn't say. If the best the hardware can technically do is 1.3M req/s, then you probably can call it a day. But if the best the hardware can do is technically 100M req/s, then you just went from very very bad (0.25% of hardware peak) to just very bad (1.2% of hardware peak). Knowing how many reqs per second should the hardware be able to do is the only way to put things in perspective here.
- slver 5y agoTCP is not typically a hardware feature so how’d you know exactly? Maybe you wanna write a dedicated OS for it? Interesting project but I can’t blame them for not doing it.
- stingraycharles 5y agoOffloading TCP to hardware is, in fact, something that is very common, especially once you get into the 10gbit connections area. I would be surprised if AWS didn’t do this.
- slver 5y agoIt's available, is it very common, I can't claim. Googling stuff like "Amazon AWS hardware TCP TOE" doesn't reveal anything. So we can't assume that either.
- jiggawatts 5y agoTypically with public cloud vendors you get SR-IOV networking above a certain VM size, but you may have to jump through hoops to enable it. I'm not sure about AWS, but in Azure it is called "Accelerated Networking" and it is available in most recent VM sizes that have 4 CPUs or more. It enables direct hardware connectivity and all offload options. In my testing it reduces latency dramatically, with typical applications seeing a 5x faster small transactions. Similarly, you can get "wire speed" for single TCP streams without any special coding.
- HugoDaniel 5y ago"Disabling these mitigations gives us a performance boost of around 28%. " This can't be serious. Can someone flag this article? Highly inappropriate.
- pornel 5y agoInteresting that most of the gains are from better utilization/configuration of Linux, not from code optimizations. The userland code was, and remained a tiny fraction of time spent.
- ameyv 5y agoHi Marc, Fantastic work! Keep it up.
- romanitalian 5y agoDo you compare with Japronto?
- romanitalian 5y agoDo you see "Japronto" [https://github.com/squeaky-pl/japronto https://github.com/squeaky-pl/japronto] ?
- sigg3 5y agoI'm digging the website layout.What's the CSS framework he's using? I'm on mobile, and can't see the source.
- Adiqq 5y agoAnyone can recommend similar articles/blogs that focus on optimization of networking/computing in Linux/cloud environments? This kind of articles are very informative, because they refer to advanced mechanisms that I either haven't heard about or newer saw in practical use.