4 ms·
But that's a bit absurd isn't it? A recent core i7 reaches 124850 MIPS. That means that if it takes 1/1000th of a second to handle a connection, the CPU could e
by pyalot2 13y ago
But that's a bit absurd isn't it? A recent core i7 reaches 124850 MIPS. That means that if it takes 1/1000th of a second to handle a connection, the CPU could execute 124 million instructions during that time. At 10/10000th it could still execute 12 million instructions. A minimal asynchronous connection handling certainly doesn't require more than say 10'000 instructions, so our servers are under-delivering on their performance at least by a factor of 1200x, perhaps even by a factor of 12000x.
Our servers should be able to surpass C1k easily, even C10k shouldn't tax them. They should, by rights, only be taxed by the C10m problem.
- sharpneli 13y agoThere is one number that has not really changed. It's memory latency, and another is the processor clock speed. The latency of main memory read is still around 100ns. It has been around that for over 10 years now. It means your CPU will have to wait for hundreds of clock cycles to get a read from RAM if it's not in the cache, and in huge datasets it probably is not in cache. Another issue is the processor clock speed. Yes it is true that modern i7 can reach 124850 MIPS. However that number comes from having 4 cores with each of them being able to reach up to 8 instructions per clock. You are still limited in executing dependent instructions. That sounds a lot. But one must remember that it reaches 8 instructions per clock only when the instructions are a good mix of float/int instructions, no branches and the instructions are not dependent on eachother. In practice you reach maybe 1-2 instructions per clock. In some code it can go even to 0.5 IPC (bunch of unpredictable branches and whatnot). Writing a code that takes advantage of large memory bandwidth and poor latency combined with massive CPU performance if the instructions are not too dependent on eachother is almost like writing modern GPU programs. It would be interesting to see what kind of an web server perf one could get by carefully writing it in OpenCL (using CPU target, not GPU).
- pyalot2 13y agoYeah I'm not disputing that there aren't bottlenecks in the system (it's not only the memory, the bus between NIC and CPU is also to blame). But, blaming bad I/O performance on large datasets misses the point slightly. You can perfectly well write a program that doesn't use heap, and has less stack use than those processors have L2 cache... (of course that's a test program). But the network performance is probably not bound by system latency as much as by an abysmally bad software stack, started with the kernel to the networking stack to the sheer idea of TLS and to the implementation of TLS (OpenSSL). Indeed, there's been calls to get rid of it all, all the layers and whatnot, and bann the OS from all but one or two cores and get rid of the whole network stack and layers and implement the networking directly in the application that needs to do it.
- rbanffy 13y ago> there's been calls to get rid of it all, all the layers and whatnot, and bann the OS from all but one or two cores and get rid of the whole network stack and layers and implement the networking directly in the application that needs to do it You certainly know building and supporting that would cost more or less the same as building and operating a sizeable datacenter. If it succeeds. Using all processing power a modern CPU offers on real code with real data is almost impossible. And it's not only memory latency and instruction interdependence - there are latencies all over a PC even before you leave the rackmount chassis. The supporting network is another source of uncontrollable latencies. Most apps I manage spend 99.99% of their time waiting for something to happen, be it the next packet, be the results from another server, which is actually a cluster behind one or more load balancers. You may get some better cache hit ratios by tweaking thread/core affinity, but it won't take you to where you want to be. If you really need that much performance, I'd suggest building your own VLIW architecture and generate the instruction mix on auxiliary CPUs as a single continuous thread on the fly based on all incoming requests for the VLIW core to devour. That would be a huge undertaking, but it would also be pretty cool CompSci.
- vidarh 13y ago> You certainly know building and supporting that would cost more or less the same as building and operating a sizeable datacenter. If it succeeds. There are plenty of solutions for that already. For the simplest case of the user-space networking, you can pick a number of "off the shelf" solutions for it: http://lukego.github.io/blog/2013/01/04/kernel-bypass-networking/ http://lukego.github.io/blog/2013/01/04/kernel-bypass-networ... http://www.openonload.org/ http://www.openonload.org/