2 ms·
Yeah I'm not disputing that there aren't bottlenecks in the system (it's not only the memory, the bus between NIC and CPU is also to blame). But, blaming bad I
by pyalot2 13y ago
Yeah I'm not disputing that there aren't bottlenecks in the system (it's not only the memory, the bus between NIC and CPU is also to blame).
But, blaming bad I/O performance on large datasets misses the point slightly. You can perfectly well write a program that doesn't use heap, and has less stack use than those processors have L2 cache... (of course that's a test program).
But the network performance is probably not bound by system latency as much as by an abysmally bad software stack, started with the kernel to the networking stack to the sheer idea of TLS and to the implementation of TLS (OpenSSL).
Indeed, there's been calls to get rid of it all, all the layers and whatnot, and bann the OS from all but one or two cores and get rid of the whole network stack and layers and implement the networking directly in the application that needs to do it.
- rbanffy 13y ago> there's been calls to get rid of it all, all the layers and whatnot, and bann the OS from all but one or two cores and get rid of the whole network stack and layers and implement the networking directly in the application that needs to do it You certainly know building and supporting that would cost more or less the same as building and operating a sizeable datacenter. If it succeeds. Using all processing power a modern CPU offers on real code with real data is almost impossible. And it's not only memory latency and instruction interdependence - there are latencies all over a PC even before you leave the rackmount chassis. The supporting network is another source of uncontrollable latencies. Most apps I manage spend 99.99% of their time waiting for something to happen, be it the next packet, be the results from another server, which is actually a cluster behind one or more load balancers. You may get some better cache hit ratios by tweaking thread/core affinity, but it won't take you to where you want to be. If you really need that much performance, I'd suggest building your own VLIW architecture and generate the instruction mix on auxiliary CPUs as a single continuous thread on the fly based on all incoming requests for the VLIW core to devour. That would be a huge undertaking, but it would also be pretty cool CompSci.
- vidarh 13y ago> You certainly know building and supporting that would cost more or less the same as building and operating a sizeable datacenter. If it succeeds. There are plenty of solutions for that already. For the simplest case of the user-space networking, you can pick a number of "off the shelf" solutions for it: http://lukego.github.io/blog/2013/01/04/kernel-bypass-networking/ http://lukego.github.io/blog/2013/01/04/kernel-bypass-networ... http://www.openonload.org/ http://www.openonload.org/