4 ms·
"The ruby version is single threaded and the test is on a 32 core workstation (counting CPU time Python is only 4x faster and OCaml is only 17x faster)" I ment
by mfp 18y ago
"The ruby version is single threaded and the test is on a 32 core workstation (counting CPU time Python is only 4x faster and OCaml is only 17x faster)"
I mentioned that on my blog. I also have an OCaml version that is 25x faster in CPU time (i.e., as fast as the top C++ entries), but barely faster regarding wall clock time. It takes a couple dozen extra lines.
Keep in mind that the Wide Finder 2 benchmark was about parallelism from the beginning; I said the Ruby version was naïve precisely because it wasn't parallel. The fact that the language did matter to this extent came as a relative surprise, because the most expensive operations in the Ruby version actually take place in its core classes, written in C. It's just that it's so slow everywhere else that the overall performance is still an order of magnitude worse.
"The author admits that there were multiple stabs at the OCaml version. What savings could come from optimizing the code in other languages?"
There are three OCaml versions, listed on the result table
http://wikis.sun.com/display/WideFinder/Results http://wikis.sun.com/display/WideFinder/Results
AFAIK other entries received considerable optimization effort (I'd even go further and say that most involved more) --- several went through half a dozen revisions, even if the wiki doesn't reflect it.
You can take a look at the wide-finder mailing list to see how often each participant was using the T2K (we used the ML to reserve time slots): http://groups.google.com/group/wide-finder/topics?hl=en&start=160&sa=N http://groups.google.com/group/wide-finder/topics?hl=en&...
wf2_multicore.ml was the first version I ran against the full dataset on the T2K, and did quite well (8 minutes). The 2(?) first runs crashed because I exhausted the memory space of the T2K, but the 3rd one completed successfully.
wf2_multicore2_block.ml took considerably more time because I switched from line-oriented to block-based IO --- basically the technique all the fast implementations used.