6 ms·
With nginx and 256 core Epycs, most single servers can easily do 200k requests per sec. Very few companies have more needs
by trueismywork 10mo ago
With nginx and 256 core Epycs, most single servers can easily do 200k requests per sec. Very few companies have more needs
- intothemild 10mo agoI can't tell if this is sarcasm or not. They didn't have this kind of compute back when the article was written. Which is the point in the article.
- trueismywork 10mo agoHalf serious. I guess what Iwas saying is that it is that kind of science which is still very useful but more to nginx developers themselves. And most users now dont have to worry about this anymore. Should have prefixed my comment wirh "nowadays"
- hinkley 10mo agoIn spring 2005 Azul introduced a 24 core machine tuned for Java. A couple years later they were at 48 and then jumped to an obscene 768 cores which seemed like such an imaginary number at the time that small companies didn’t really poke them to see what the prices were like. Like it was a typo.
- fweimer 10mo agoBefore clusters with fast interconnects were a thing, there were quite a few systems that had more than a thousand hardware threads: https://linuxdevices.org/worlds-largest-single-kernel-linux-system/ https://linuxdevices.org/worlds-largest-single-kernel-linux-... We're slowly getting back to similarly-sized systems. IBM now has POWER systems with more than 1,500 threads (although I assume those are SMT8 configurations). This is a bit annoying because too many programs assume that the CPU mask fits into 128 bytes, which limits the CPU (hardware thread) count to 1,024. We fixed a few of these bugs twenty years ago, but as these systems fell out of use, similar problems are back.
- alexjplant 10mo ago> Driven by 1,024 Dual-Core Intel Itanium 2 processors, the new system will generate 13.1 TFLOPs (Teraflops, or trillions of calculations per second) of compute power. This is equal to the combined single precision GPU and CPU horsepower of a modern MacBook [1]. Really makes you think about how resource-intensive even the simplest of modern software is... [1] https://www.cpu-monkey.com/en/igpu-apple_m4_10_core https://www.cpu-monkey.com/en/igpu-apple_m4_10_core
- fweimer 10mo agoNote that those 13.1 TFLOPs are FP64, which isn't supported natively on the MacBook GPU. On the other hand, local/per-node memory bandwidth is significantly higher on the MacBook. (Apparently, SGI Altix only had 8.5 to 12.8 GB/s.) Total memory bandwidth on larger Altix systems was of course much higher due to the ridiculous node count. Access to remote memory on other nodes could be quite slow because it had to go through multiple router hops.
- hinkley 10mo agoMy Apple Watch can blow the doors off a Cray 1. It’s crazy.
- marcosdumay 10mo agoThe article was written exactly because they had machines capable enough at the time. But the software worked against it on every level.
- yencabulator 9mo agoI mean, yes and no. It was a software challenge to hit the hardware limit, but the hardware limits were also much lower. My team stopped optimizing when we maxed out the PCI bus in ~2001.
- Maxatar 10mo agoI don't see how you could have read the article and come to this conclusion. The first few sentences of the article even go into detail about how a cheap $1200 consumer grade computer should be able to handle 10,000 concurrent connections with ease. It's literally the entire focus of the second paragraph. 2003 might seem like ancient history, but computers back then absolutely could handle 10,000 concurrent connections.
- api 10mo agoI’m shocked that a 256 core Epyc can’t do millions of requests per second at a minimum. Is it limited by the net connection or is there still this much inefficiency?
- otterdude 10mo ago256 Processes x 10k clients (per the article) = 256K RPS
- mrweasel 10mo agoAren't you of by a zero? 10K requests per core / per second, time 256 cores is 2.560.000 RPS. There's probably going to be some overhead, but it seems like you could do 1M, if you have the bandwidth.
- cap11235 10mo agoI think the most likely bottleneck is gonna be your NIC hating getting a ton of packets. Line rate with huge frames is quite different than line rate with just ICMP packets, for instance (see CME binary glink market data for a similarly stressful experience to the ICMP).
- tempest_ 10mo agoLike anything it really depends on what they are doing, if you wanted to just open and close a connection you might run into bottle necks in other parts of the stack before the CPU tops out but the real point is that yea, a single machine is going to be enough.
- zipy124 10mo agoIt almost certainly can, even old intel systems with dual CPU 16 core systems could do 4 and a half million a second [1]. At a certain point network/kernel bottlenecks become apparent though, rather than being compute limited. [1]: https://www.amd.com/content/dam/amd/en/documents/products/ethernet-adapters/onload/onload-nginx-plus-benchmark-results.pdf https://www.amd.com/content/dam/amd/en/documents/products/et...
- tempest_ 10mo agoThis is how I feel about this industries fetishization of "scalability". A lot of software time is spent making something scalable when in 2025 I can probably run any site the bottom 99% of most visited sites on the internet on a couple machines and < 40k capital.
- tbrownaw 10mo ago> any site the bottom 99% of most visited sites on the internet What % is the AWS console, and what counts as "running" it?
- tempest_ 10mo ago> What % is the AWS console 0% Prior to the recent RAM insanity(a big caveat I know) a 1u supermicro machine with 768GB some NVME storage and twin 32 core Epyc 9004s was ~12K USD. You can get 3 of those and and some redundant 10G network infra(people are literally throwing this out) for < 40k. Then you just have to find a rack/internet connection to put them in which would be a few hundred a month. The reality is most sites don't need multi region setups, they have very predicable load and 3 of those machines would be massive overkill for many. A lot of people like to think they will lose millions per second of down time, and some sites certainly do but most wont. All of this of course would be using new stuff. If you wanted to use used stuff the most cost effective are the 5 year old second gen xeon scalables that are being dumped by cloud providers. Those are more than enough compute for most they are just really thirsty so you will pay with the power bill. This of course is predicated on assumption you have the skill set to support these machines and that is increasingly becoming less common though as successful companies that started in the last 10 years are starting to do more "hybrid cloud" it is starting to come back around.
- cap11235 10mo agoIf you are paying 12k, why would you ever subject yourself to supermicro
- 9mo ago
- otterdude 10mo agoWhen people talk about a single server they're not talking about one hunk of metal, they're talking about 1 server process. This article describes the 10k client connection problem, you should be handling 256K clients :)
- marcosdumay 10mo agoWhen people talk about a single server they are pretty much talking about either a single physical box with a CPU inside or a VPS using a few processor threads. When they say "most companies can run in a single server, but do backups" they usually mean the physical kind.
- Maxatar 10mo agoThe term is absolutely ambiguous and I know I've run into confusion in my own work due to the ambiguity. For the purpose of the C10K, server is intended to mean server process rather than hardware.
- marcosdumay 10mo ago> You can buy a 1000MHz machine with 2 gigabytes of RAM and an 1000Mbit/sec Ethernet card for $1200 or so. Let's see - at 20000 clients, that's 50KHz, 100Kbytes, and 50Kbits/sec per client. It shouldn't take any more horsepower than that to take four kilobytes from the disk and send them to the network once a second for each of twenty thousand clients. It was about physical servers.
- Aloisius 10mo agoThis is definitely talking about scaling past 10K open connections on a single server daemon (hence the reference to a web server and an ftp server). However, most people used dedicated machines when this was written, so scaling 10K open connections on a daemon was essentially the same thing as 10K open connections on a single machine.
- dilyevsky 10mo agoAt the time this was written powerful backend server only had like 4 cores. Linux only started adopting SMP like that same year. Also CPU caches were tiny Serving less than 1k qps per core is pretty underwhelming today, at such a high core count you'd likely hit OS limitations way before you're bound by hardware
- Grosvenor 10mo agoLinux had been doing SMP for about 5 years by that point. But you're right OS resource limitations (file handles, PIDs, etc) would be the real pain for you. One problem after another. Now, the real question is do you want to spend your engineering time on that? A small cluster running erlang is probably better than a tiny number of finely tuned race-car boxen.
- dilyevsky 10mo agoMy recollection is fuzzy but i remember having to recompile 2.4-ish kernels to enable SMP back in the day which took hours... And I think it was buggy too. Totally agree on many smaller boxes vs bigger box especially for proxying usecase.