6 ms·
Performance of Python runtimes on a non-numeric scientific code
- mih 12y agoTL;DR - Nuitka (which has been garnering some attention on HN today - https://news.ycombinator.com/item?id=8771925 https://news.ycombinator.com/item?id=8771925) is not as fast as CPython and less faster than PyPy for the chosen scenario (enumeration of a fat graph and computing their graph homology). A year and a half down the line after this paper was published (Apr 2013), Nuitka's main strengths seem to be the convenience in packaging the code into an executable which a number of readers have reported success in. If runtime and memory usage are more important, you are better off sticking with the interpreters for now.
- mkesper 12y agoThis should be re-done with up to date versions of all programs. The authors mention having opened bug reports/existing known bugs. All contestants should have time to react to those, shouldn't they?
- reikonomusha 12y agoNot necessarily. A sort of test with a somewhat random computation where implementations have not specifically tuned for said test can provide beneficial information. If you have enough of such tests, one can get a broad idea of the performance of such systems. The worst case is when people hyperoptimize for a particular benchmark, which tells you nothing. You see this especially in the Benchmarks Game.
- andreasvc 12y agoWhile that would be interesting, I don't see it as the responsibility of the author of a conference poster that has already been presented. It would make more sense for the authors of the various Python implementations to take this up as a benchmark.
- reikonomusha 12y agoI didn't find this paper to be very good. While it talked lightly about a relatively complex mathematical object it computes, it did not talk much about what's involved in its computations, except for some very high level keywords ("comprehensions", "object graphs"). What algorithms were used? What data structures? Was the code idiomatic? Was there any effort to reduce things like allocation? Was homological computation the only test case? Even numerical benchmarks typically come in a suite (a good sprinkling of linear algebraic computations, tight straight line floating point programs, differential equation solvers, various numerical simulators, ...), because one LAPACK function will not give you the full picture. This paper did not give me a very good understanding of how performant non-numeric math—which in and of itself is an extremely broad and general term—is on each implementation.
- dalke 12y agoIt's a conference paper, from EuroSciPy 2013, distributed through a preprint service. There's no expectation it will be a high quality paper. Instead, it's an appropriate quality for where and how it was published. "What algorithms were used"? The papers says "The code used to install the software and run the experiments is available on GitHub at https://github.com/riccardomurri/python-runtimes-shootout https://github.com/riccardomurri/python-runtimes-shootout " Checking it now, it gives a reproducible way to download the specific packages used, and the benchmark framework. The actual code benchmarked is fatghol, from https://code.google.com/p/fatghol/ https://code.google.com/p/fatghol/ . There's also a link to a preprint describing the construction algorithm, at http://arxiv.org/pdf/1202.1820v2.pdf http://arxiv.org/pdf/1202.1820v2.pdf . What you propose is an unrealistic expectation, and only possible for people with lots of money and time. Instead, in real life what happens is people do A, and publish A, then do B (building on A), and publish B, then do C (building on B) and publish C. There's a trail of work backing up the final publication. It makes no sense for publication Q to revisit all of A-P, nor for the author to wait until Z before finally publishing everything. I also think knowledge transfer would be lower since someone interested in this paper's conclusions about the available documentation for the different Pythons (EuroSciPy is not a graph theory specialist conference) would almost certainly not be interested in the algorithm generation details. You do realize the LINPACK is the "gold standard" benchmark used to rank the top 500 supercomputers, right? And all it does is solve A x = B. In any case, the performance suites like SPEC MPI still need to evaluate the individual benchmarks before assembling them in a suite. Even if you require a suite for something to be meaningful to you, this could be seen as a first step to building such a meaningful suite. It appears to me, therefore, like you are needless harsh and critical.
- TazeTSchnitzel 12y agoBy all appearances, the alternatives are only negligibly faster than CPython?!
- chrisseaton 12y agoDid you see that those graphs are logarithmic? PyPy looks between 3x and 4x on problems that run long enough for PyPy to have a chance to JIT properly.
- andreasvc 12y agoThis comparison is interesting in that it compares the performance of plain or type-annotated Python code, but to get the full performance benefit of Cython you would replace lists with C arrays, objects with C structs, etc.