6 ms·
The 10k problem is still not solved. A run of the mill VPS typically manages around 1k simultaneous connections and 10k requests/s. That is, over plain http. Ov
by pyalot2 13y ago
The 10k problem is still not solved. A run of the mill VPS typically manages around 1k simultaneous connections and 10k requests/s. That is, over plain http. Over TLS this goes down to about 100 connections and 600 requests/s. A cheap amazon instance will usually be around 1/5th of all that.
With a bit of tweaking you can get the plain http case a bit higher, but the TLS route will not get much better, because TLS isn't written with speed in mind (the handshake is slow and expensive) and the prevalent implementation (OpenSSL) is not written for high performance servers (it basically dictates that you run blocking sockets in a thread per connection).
Unfortunately spdy and various http2 proposals rely on TLS (in order to punch trough proxies), which means going back in server performance about 10 years.
So it is of little surprise that companies have started offering "cloud" solutions, because the typical VPS can't handle todays high traffic internet over TLS (worse than plain http by a factor of 10) and the typical cloud server is worse than a VPS by a factor 5, creating a 50x performance degradation, artificially. Obviously when faced with the question of running 10 servers, or 100, most small companies turn to the even worse "cloud" solution, requiring even more servers (500).
The whole affair is a sodding mess, and we're wasting massive amounts of energy and capital on insisting on doing things inefficiently. This is because by rights our VPS servers should easily be able to break trough the 10k limit in every way, but it can't because the OS wastes a lot of time running an inefficient network stack as well as that TLS and OpenSSL can't be bothered to get their act together.
And that is how, in the year of our lord 2014, more than 15 years after somebody writing the 10k problem article, and after webservers becoming at least 128x faster then back in the 90ties, most websites out there can still not stand up to serious traffic, take ages to load, and are hosted by an infrastructure (routers and whatnot) that easily buckle under even light DDOSing.
- wbsun 13y agoWhen it goes to transactions and stateful sessions, C10K, or even C1K is far from solved yet.
- pyalot2 13y agoBut that's a bit absurd isn't it? A recent core i7 reaches 124850 MIPS. That means that if it takes 1/1000th of a second to handle a connection, the CPU could execute 124 million instructions during that time. At 10/10000th it could still execute 12 million instructions. A minimal asynchronous connection handling certainly doesn't require more than say 10'000 instructions, so our servers are under-delivering on their performance at least by a factor of 1200x, perhaps even by a factor of 12000x. Our servers should be able to surpass C1k easily, even C10k shouldn't tax them. They should, by rights, only be taxed by the C10m problem.
- sharpneli 13y agoThere is one number that has not really changed. It's memory latency, and another is the processor clock speed. The latency of main memory read is still around 100ns. It has been around that for over 10 years now. It means your CPU will have to wait for hundreds of clock cycles to get a read from RAM if it's not in the cache, and in huge datasets it probably is not in cache. Another issue is the processor clock speed. Yes it is true that modern i7 can reach 124850 MIPS. However that number comes from having 4 cores with each of them being able to reach up to 8 instructions per clock. You are still limited in executing dependent instructions. That sounds a lot. But one must remember that it reaches 8 instructions per clock only when the instructions are a good mix of float/int instructions, no branches and the instructions are not dependent on eachother. In practice you reach maybe 1-2 instructions per clock. In some code it can go even to 0.5 IPC (bunch of unpredictable branches and whatnot). Writing a code that takes advantage of large memory bandwidth and poor latency combined with massive CPU performance if the instructions are not too dependent on eachother is almost like writing modern GPU programs. It would be interesting to see what kind of an web server perf one could get by carefully writing it in OpenCL (using CPU target, not GPU).
- pyalot2 13y agoYeah I'm not disputing that there aren't bottlenecks in the system (it's not only the memory, the bus between NIC and CPU is also to blame). But, blaming bad I/O performance on large datasets misses the point slightly. You can perfectly well write a program that doesn't use heap, and has less stack use than those processors have L2 cache... (of course that's a test program). But the network performance is probably not bound by system latency as much as by an abysmally bad software stack, started with the kernel to the networking stack to the sheer idea of TLS and to the implementation of TLS (OpenSSL). Indeed, there's been calls to get rid of it all, all the layers and whatnot, and bann the OS from all but one or two cores and get rid of the whole network stack and layers and implement the networking directly in the application that needs to do it.
- Nursie 13y agoThere are many commercial alternatives to OpenSSL. I've worked with one or two. There's bound to be one that suits your needs. Of course you may also need to go to a full commercial, non-FOSS stack, but I'm sure plenty of companies will have those to sell too.
- theknown99 13y agoNot really true unless you use crappy software. I run a few thousand simultaneous connections on AWS small instances without any issues. (HTTP + HTTPS). The key is writing good software. Which a lot of programmers are still really bad at.
- pyalot2 13y agoYou know, it's cute when somebody who wasn't even born when I started writing software for servers tries to discredit me.
- rbanffy 13y agoIt would be better if you both gave more details on what and how you do so the possible mistakes from both sides could be pointed out and we could learn from them. From a quick glance on your post, I too suspect there is something wrong - your performance shouldn't be that bad, but there is not enough information to point where.
- pyalot2 13y agoThat's really too long to list, I could conceivably write a book about it. But, fortunately it's relatively easy to test. You get whatever server you prefer, and install nginx and install a bunch of performance testing tools (like siege, ab etc.) and then you test different concurrency load scenarios. No custom software, and widely regarded the fastest webserver out there.
- rbanffy 13y agoAs I pointed elsewhere, most recent servers are just overgrown IBM 5150 PCs, with much faster CPUs, immense amounts of memory and storage and somewhat faster buses. They are desktop PCs misused as servers.
- theknown99 13y ago> and install nginx Yeah there's your problem...
- seunosewa 13y agoIf you care about efficiency you should not use VPSs. You should dedicated servers.