5 ms·
C10K 2012: Erlang Wins, Go a close second. Java, Haskell, Node & Python fail
- KirinDave 14y agoReally. Bad. Title.
- simfoo 14y agoTry again and increase GOMAXPROC to 10-20. See for example here: https://groups.google.com/d/msg/golang-nuts/fjBft6qeMo0/kYSiFTX2j1sJ https://groups.google.com/d/msg/golang-nuts/fjBft6qeMo0/kYSi...
- adamtulinius 14y agoThe numbers are confusing and doesn't seem to add up at all. The section "stat definitions" describes data not available in the table below.
- KirinDave 14y agoPart of the problem is that EC2 is such a whacky environment. Notice Erlang's median connection time was tiny but its average had huge outliers. Read the timing numbers with a reasonable portion of salt. It's EC2 we're talking about here.
- tedsuo 14y agoYeah I wince a bit at any benchmarking being run in a shared environment like EC2, where other activity outside of your VM can affect performance.
- StavrosK 14y agoHmm, why would he run a benchmark on a shared environment? Didn't he have a home computer to run this on?
- wmf 14y agoWhich is more repeatable but less representative. IMO a non-shared EC2 instance is probably the way to benchmark (if you can drive enough load to saturate it).
- StavrosK 14y agoI agree, but does EC2 have non-shared instances?
- gojomo 14y agoHmm. Some cloud provider ought to provide (for a premium price) guaranteed-identical non-shared configurations... including small groups of machines with uncontended, identical cross-connects. (I know it's then close to 'dedicated' hosting, but people running benchmarks also want the quick setup and discard of virtualized instances. This would be a hybrid offering that ensures their cloud services is always chosen for such benchmarking comparisons. Of course this offering would not be good for cross-cloud comparisons, because it's not representative of their usual offerings.)
- wmf 14y agoguaranteed-identical non-shared configurations... including small groups of machines with uncontended, identical cross-connects. That's called EC2 cluster compute.
- bartman 14y agoOn EC2 there are two ways to get a dedicated machine: - Using VPC you can specify that your instance should run on dedicated hardware [1] - Cluster Compute instances [2] most likely (see [3]) run on dedicated hardware too [1] http://aws.amazon.com/dedicated-instances/ http://aws.amazon.com/dedicated-instances/ [2] http://aws.amazon.com/ec2/#instance http://aws.amazon.com/ec2/#instance [3] https://forums.aws.amazon.com/message.jspa?messageID=238197#238197 https://forums.aws.amazon.com/message.jspa?messageID=238197#...
- j2labs 14y agoI totally agree, but I'm not sure of any other places that offer such flexibility in pricing and hardware. If you were to conduct a test like this, where would you go for more reliable performance from hardware?
- KirinDave 14y agoI think EC2 is great. It's just whacky.
- deleted 14y ago[deleted]
- caf 14y agoThis is true, but it's still interesting if you're intending to deploy onto EC2.
- bcx 14y agoI agree, this is really a test of various websocket implementations. (which is still cool)
- Weltschmerz 14y agoAny reason you haven't tested the 'ws' node version I submitted? I think it should work fine...
- stock_toaster 14y agoThis appears to be pointing to the old results. I dont think the author of the benchmark has rerun them yet.
- nirvana 14y agoI'm not the author of the benchmarks. I merely submitted the results to hacker news because I found them interesting.
- KirinDave 14y agoFor Haskell, there is no doubt it could do better. The culprit is probably the relatively new Websocket server. For evidence, I cite http://www.yesodweb.com/blog/2011/03/preliminary-warp-cross-language-benchmarks http://www.yesodweb.com/blog/2011/03/preliminary-warp-cross-... But "Not bad for 3 lines of code" is a pretty fair sentiment. Erlang and Go doing great is no surprise. Java didn't do too well, but the implementation didn't really use the most performant tools available. The real loser here is Node.js. Nearly as long as the Java example, but the worst performer in the real metrics. So much for "making concurrency easy."
- mrj 14y agoThis doesn't seem to investigate tuning at all. I see the Tornado script forks processes, for example, but it doesn't experiment with different values (or PyPy?). The Java code appears to be run with default VM settings, which pretty much always deserves some tweaks to run in a server environment. Those are just the two I'm most familiar with. The rest seem to have similar faults. Plus, it would be a more interesting test if the server had to perform some kind of work. Simply echoing the request is not a typical usage and could really bias the results in favor of setups that would fail on a real project. I better title might be, "Erlang wins unrealistic test over other VMs in their default configuration."
- dvirsky 14y agoI contributed the tornado script to this project, but the results of it are yet to be published (or if they have I've missed them). I haven't tested it myself on AWS, just on a physical desktop machine, I got great results, but they are so good it of course seems to be an irrelevant comparison. I have benchmarked the ws4py code in the test on that machine, and it was much slower than the tornado code - quite obvious since it ran in a single process thus on a single CPU. BTW The default number of processes when you fork is the number of CPUs, which sounded like a fair estimate to me, not knowing the target machine, so I left it.
- huggyface 14y agoThis doesn't seem to investigate tuning at all. This is always the response to any benchmark where one's pet technologies don't win. If you have a magic quadrant of tuning, put it forth. If not, you have said nothing that counters the results.
- hermanhermitage 14y agoIndeed. It seems pretty clear to me it is framed as an benchmark of non tuned idiomatic performance. Maybe it doesn't call it out explicitly - but in my view that is a reasonable enough test. I'm pretty sure all the platforms could support an FFI binding or equivalent to an optimized epoll C implementation - or hours of tuning.
- 14y ago
- rvirding 14y agoI am not competent to judge the tests Eric used. But the obvious reply to those who complain about his test for some language is to fix it and come with an improvement which you feel better represents that language. It's all on github so fork it and come with a pull-request. To say it bluntly, "put your money where your mouth is".
- eps 14y agoIs this a joke? Where is a simple epoll-based C server for a baseline comparision? :)
- anacrolix 14y agoI agree. When it outperforms the rest by a factor of 10-100 to 1, we can all laugh at the irrelevance of benchmarks.
- willvarfar 14y agoyeah, I'd love to see him enter a hellepoll-based one
- lnanek2 14y agoTitle doesn't agree with the article...article says Java beats Go, orders Java above it in the ranking chart, and the raw numbers say Java dropped fewer connections...
- tuxychandru 14y agoBut it returned only half as many messages as Go.
- dchest 14y agoDiscussed 3 days ago http://news.ycombinator.com/item?id=4105317 http://news.ycombinator.com/item?id=4105317
- jrockway 14y agoMy conclusion is that he managed to write the fastest code in the language he works with most.
- deleted 14y ago[deleted]
- kogir 14y agoBenchmarking anything on EC2 instances will yield statistically suspect results. You know nothing about the other workloads on the box, the underlying hardware, or the network connectivity.
- deleted 14y ago[deleted]
- gorset 14y agoI think he risks dropped connections because the listen queue can overflow. He doesn't mention increasing somaxconn, which means it probably has a default of 128 (I don't think increasing tcp_max_syn_backlog is needed since he's using syncookies). When creating a new connection every 1ms it only takes 128ms for the backlog to fill up, which is not that improbable with GC and JIT pauses. Java should be faster than erlang most of the time, but the pauses can kill you if you don't handle them gracefully. If the benchmark had been about "normal" http requests, I would suggest putting HAProxy in front with a reasonable maxconn - then HAProxy will hold on to the connections until the application starts accepting again.
- invisible 14y agoThis test is lacking some due diligence that I feel is skewing the results significantly (and, I assume, there are more that I haven't even noticed in Python/Java). (While m1.medium only has one CPU, the network IO cost exists.) Erlang is automatically threaded by nature so it has some inherit scaling built-in (so the code benefits from threading as long as it runs correctly). The Go code is set up to spin up threads inside of ListenAndServe thus gaining the benefit of splitting up IO. The Haskell code has a thread specifically for garbage collection, thus utilizing the second core. The Node.js code could be using cluster (or threads/fibers more directly) but isn't for some reason. It also seems this was using the websocket npm (some unknown code running on the stack!). For a valid test the websocket code should be written in JavaScript directly in the test itself. Edit: Researched and Go actually uses ListenAndServe which creates threads on-the-fly but the m1.medium is bound to 1 CPU (it still benefits from threading due to IO).
- stock_toaster 14y ago> The Go code is set to spin up threads to accomodate the number of CPUs available on the system (2 logical cores on m1.medium). m1.medium is a single virtual cpu (it is a vm, so no 'cores' to speak of). c1.medium has 2.
- invisible 14y agoCorrected in my post to reflect that - it was kind of difficult to find a real answer when I was looking that up.
- deleted 14y ago[deleted]
- DrJosiah 14y agoThere are at least 2 better http servers for Python, uWSGI and gEvent. Tornado is known to be slower: http://nichol.as/benchmark-of-python-web-servers http://nichol.as/benchmark-of-python-web-servers . More specifically, uWSGI has been shown to respond at under 25ms with 15k+ concurrent connections. Throw uWSGI behind Nginx (with it's new websocket support), tune it a bit, and I wouldn't be surprised to see it "pass" and perhaps even be competitive.
- keymone 14y agopeople get so much butthurt when they don't see "<stuff i chose to work with> wins this benchmark" any framework, any language, hell any hardware is either generalized to deal with many problems or specifically designed to battle one. you sure can write fast C program that does exactly this benchmark well - what will that prove? nothing at all.. why is it so hard to accept the fact that some tools can provide good enough results without too much tweaking while other tools may provide better result with more time spent achieving it?
- mokus 14y agoCould anyone explain to me why some of the rows don't add up to 10k attempted connections? For example, looking at the raw data for Haskell and java, and adding up connections, disconnects, crashes, and timeouts doesn't give anywhere near 10k. What happened to the rest of the connections? Were they simply not attempted? Was the port closed? Or am i just missing a relevant field in the output? It's not clear to me from the description.