3 ms·
EDIT. OK, so. I ran this benchmark myself on an 8-core Xeon running Linux. 2.13 GHz, CentOS 6.2. Kernel was 2.6.32-220.el6.x86_64. 50 gigs of RAM. I got s
by cmccabe 13y ago
EDIT. OK, so. I ran this benchmark myself on an 8-core Xeon running Linux. 2.13 GHz, CentOS 6.2. Kernel was 2.6.32-220.el6.x86_64. 50 gigs of RAM.
I got somewhere between 2.0 and 2.2 "milliseconds per ping" for Scala 2.9, and somewhere between 3.5 and 3.7 for Go 1.1. This is not the 10x difference that the authors reported, but it is something. The difference may be due in part to the different platform and hardware I am using.
Contrary to what I wrote earlier, I noticed that GOMAXPROCS=8 did seem to be slower than GOMAXPROCS=4 here. I got around 4 "milliseconds per ping" with GOMAXPROCS=8. Using a mutex and explicit condition variable shaved off maybe 0.2 milliseconds on average (very rough estimation).
Again contrary to what I wrote earlier, Nagle on versus off didn't seem to matter in the Go code. I still think you should always have it off for a test like this, but on my setup I did not see a difference.
I still don't think this benchmark is showing what they think it is. I have a hunch that this is more of a scheduler benchmark than a TCP benchmark at all. I think I'd have to haul out vtune to get any further, and I'm getting kind of tired (after midnight here).
- BarkMore 13y agoNagle's algorithm is disabled by default in Go (http://golang.org/pkg/net/#TCPConn.SetNoDelay http://golang.org/pkg/net/#TCPConn.SetNoDelay).
- jongraehl 13y ago> First of all, they're testing on MacOS, which is not going to be the platform they're actually using for the server code. The backends can be very different. Unless you have evidence that Go has MacOS-only bugs, this is meaningless speculation (though I agree that it's possible that there's no problem on Linux and that people don't generally run mac servers, I'm not sure why we should privilege your hypothesis). Agreed about GOMAXPROCS=4 - that seemed questionable to me (I don't know what it does, precisely, but I don't see a 4-anything limit in the Scala code).
- jongraehl 13y agoThanks to parent (after edit) for really testing. It turns out that he was right that Go on Mac is substantially worse than Linux - my bad. Maybe the Go lib authors didn't put much effort into reading the subtly different BSD/Darwin vs Linux syscall semantics. To explain the Nagle algorithm's irrelevance to this case, we have to understand how it works. It doesn't delay any writes once you read. It only affects two consecutive small writes (my memory was fuzzy so I checked http://en.wikipedia.org/wiki/Nagle's_algorithm http://en.wikipedia.org/wiki/Nagle's_algorithm ). Odds of preemption between the client's write and its read seem small, so it shouldn't matter whether you Nagle or not.
- cmccabe 13y agoYeah, I always seem to forget the exact details of Nagle (probably because every project I've worked on just turns it off). Write-write-read is the killer, I guess-- just doing a small write followed by a small read, or vice versa, should not be affected by Nagle. So the results I got make sense. Re: MacOS, I do know that some of the Go developers use Macs as their primary desktops. So I don't think they neglect it, but given that they're targeting the server space, it makes sense to optimize Linux more. I still haven't seen any really good explanation of these results. I don't buy the argument that the JVM is providing the advantage here. The main thing that the JVM is able to do is dynamically recompile code, and this shouldn't be a CPU-bound task.
- diakritikal 13y agoIt's also hard to accurately profile Go programs on OS X because of bugs in it's now quite stale kernel. Specifically SIGPROF on OS X isn't always sent to the currently executing thread. Afaik this isn't a problem on newer FreeBSD kernels.
- trailfox 13y ago> I got somewhere between 2.0 and 2.2 "milliseconds per ping" for Scala 2.9 The latest Scala is 2.10, which has a number of optimizations over 2.9...