8 ms·
Pythran: Crossing the Python Frontier [pdf]
- accurrent 8y agoLink to the actual software: https://github.com/serge-sans-paille/pythran https://github.com/serge-sans-paille/pythran
- emmelaich 8y ago> "Pythran is an ahead of time compiler for a subset of the Python language, with a focus on scientific computing." and > a claimless python to c++ converter
- serge-ss-paille 8y ago(main dev writting) Don't focus too much on the claimless :-) There still are far more wide spread tools that solve the same kind of issues (numba, cython, julia)... The idea of the change was that it's more important to convey the problem it solves rather than how it's done ;-)
- UncleEntity 8y agoNice project, I've wanted to do something similar using my libjit python bindings but never seem to work up the gumption. Simple and Effective Type Check Removal through Lazy Basic Block Versioning[0] could be adapted to make it unnecessary to compile more than one version of the function at a (probably) minor cost of performance. It's geared more towards jitted code but something like it where it compiles different blocks where the types matter and falls back to interpreted code if users pass in some random types not expected just might work -- or the python interpreter throws an exception if the types don't make sense. [0]http://drops.dagstuhl.de/opus/volltexte/2015/5219/pdf/9.pdf http://drops.dagstuhl.de/opus/volltexte/2015/5219/pdf/9.pdf
- eljost 8y agoInteresting article but just skimming through it some things stand out immediately: 1.) The first snippet isn't even valid python code as floats don't have a shape attribute. s = 0. n = s.shape 2.) The inline latex math isn't rendered properly.
- kristofferc 8y agoThe first snippet also doesn't balance the parenthesis s += 100. * x[i + 1] - x[i] ** 2.) ** 2. + (1 - x[i]) ** 2
- rlayton2 8y agoI believe the input should be a numpy array of floats which has a shape attribute
- serge-ss-paille 8y ago(shameful author here) One should read def rosen_explicit_loop(x): s = 0. n = x.shape[0] for i in range(0, n - 1): s += 100. * (x[i + 1] - x[i] ** 2.) ** 2. + (1 - x[i]) ** 2 return s (edited)
- chestervonwinch 8y ago... n = x.shape[0]
- quietbritishjim 8y agoOr just len(x). This works perfectly well on numpy arrays and has the bonus that it works on regular lists/tuples of floats, so the first snippet doesn't rely on numpy.
- chestervonwinch 8y agoIt should be the shape of x (actually, the zero'th element of the shape), but this is also a tad odd because this would assume that x is a numpy array, which isn't introduced until after this 'naive pure python' code block (i.e., before numpy is even introduced in the text).
- fermigier 8y agoTalk from PyParis 2015: https://www.youtube.com/watch?v=Af8B30mXZ7E https://www.youtube.com/watch?v=Af8B30mXZ7E Slides: https://fr.slideshare.net/PoleSystematicParisRegion/track-32-serge-guelton-et-pierrick-brunet https://fr.slideshare.net/PoleSystematicParisRegion/track-32...
- icebraining 8y agoAlso from this year's FOSDEM: https://fosdem.org/2018/schedule/event/pythran/ https://fosdem.org/2018/schedule/event/pythran/
- superdimwit 8y agoSuspicious that they make brief mention of numba but don't make comparison to it in the following discussion or benchmarks..
- joshsyn 8y agoIsn't this solved by julia? I think scientific community should use a more functional language rather than language like python tbh
- montalbano 8y agoThough I'll still use Python for non-scientific programming, I've switched to Julia for all my scientific programming needs in my day job.
- KronenR 8y agoGive me an ecosystem like Python has and I won't switch anyways because for when that happens I will have more experience with Python, which is much more important.
- rfeather 8y agoCould you elaborate more on the advantages of a functional language?
- cbcoutinho 8y agoScientific work should strive to be functional by definition (identical input == identical output), so I could see why a programing language should reflect that.
- ianamartin 8y agoThat's a pretty broad statement. And it's a goal that a lot of scientific work simply doesn't allow for. Any stochastic process will have variability for a given set of inputs. I think there are more fields within the umbrella of Science where purely functional programming isn't an achievable ideal--let alone a desirable one--than there are where this is a good fit. You can make the argument that all programming should be functional on its own if you want, and I'm open to that. But this it's pretty sketchy to me to just shoehorn all scientific work into a category of work that should definitionally be functional. There are huge numbers of scientists who do not agree with that at all.
- HerrMonnezza 8y agoDoes anyone know how this compares to existing Python-to-C++ transpilers like Cython or Shedskin?
- Arkanosis 8y agoCython is a bit different from CPython / Pythran / Shedskin in that you need to learn the Cython language, which is a Python-ish programming language, but not Python. Shedskin and Pythran look somewhat similar to me (disclaimer: I've contributed quite a bit to Shedskin but have never used Pythran so far), except in Shedskin you don't even need annotations like with Pythran (the downside being the finer control you have of the native types used is through the transpiler options). Also, Shedskin development is not much active these days — to say the least — and there's zero support for Python 3, while Pythran is under fairly active development has beta support for Python 3. If you're interested in Python / native implementations, you might be interested by Nuitka as well: http://nuitka.net/ http://nuitka.net/
- xapata 8y ago> Python-ish It's a little more Python-like than just -ish.
- HerrMonnezza 8y agoThanks for your answer! However, I think this might be a bit misleading for people who do not know Cython: > the Cython language, which is a Python-ish programming language, but not Python. Actually, http://cython.org/ http://cython.org/ states: """ The Cython language is a superset of the Python language that additionally supports calling C functions and declaring C types on variables and class attributes. """ In my experience with Cython, this description is quite accurate: code can be annotated with C types and then compiled to efficient C code by Cython; if you don't use annotations, then you can still compile to C code but with less speed advantage. I haven't used Cython since version 0.17 (quite old now) but IIRC the major drawback was that it was mainly targeting writing extension modules for Python; it could generate self-standing executables, but would still require a Python interpreter to be embedded in any compiled code (that was the price for seamless interoperability between Cython/compiled code and "pure Python" code).
- dilawar 8y agoI wonder why the author did not compare performance of pypy? I guess pypy is jit compiler.
- dec0dedab0de 8y agoI was going to say that most scientific python libraries use numpy, but a quick google shows that pypy supports numpy now.
- AnimalMuppet 8y agoOff topic, but this seems like as good a place as any to ask: It's my impression that numpy is really good. Is it as good as Fortran? That is, if I have a large, sparse, complex matrix, Fortran will have an efficient solver for it that will also be numerically stable, and will have four decades of use to find any weaknesses. Is numpy equivalent (except for the four decades part)? Is it close? Or does it just cover the basic cases well, and for the specializations you're on your own?
- zb 8y agoNumPy is designed to work with SciPy, which is a wrapper for literally the same 4 decade old Fortran libraries (LAPACK &c.) that you're referring to here.
- gnufx 8y agoNote that LAPACK isn't four decades old (and probably should be superseded anyway). You perhaps won't want to use large-scale numerical libraries that are that old unless they've had a lot of development and been well parallelized. One issue with mixed language programming that we learned decades ago is the issue of debugging that there typically is across the interface. Also general tool support. I recently asked the local Python expert about HPC-style profiling of Python calling C(++) libraries, for instance. (I couldn't make TAU work.)
- geoalchimista 8y ago> "As a matter of comparison, Cython does not support principles 1, 2, or 3 and has optional support for 4." For 2 Type agnosticism, this can be emulated with a "fused type" in Cython. See this example: http://cython.readthedocs.io/en/latest/src/userguide/numpy_tutorial.html#more-generic-code http://cython.readthedocs.io/en/latest/src/userguide/numpy_t... But I think the major inconvenience with Cython vectorization is really not about `float32` and `float64`. You get `float64` NumPy array by default from floating-point calculations. The actual inconvenience is that the vectorized function cannot take a scalar input like the NumPy ones. To remain polymorphic, I usually have to perform an `is_scalar` check on the input in a Python wrapper before sending the input data to the Cython function.
- targafarian 8y agoNote that the alternative in Numba is incredibly convenient, where you can trivially create a function where the body of the function operates on a scalar, but then using the @vectorize decorator which makes it into a numpy ufunc automatically. This generalizes the function to operate equally well on scalars or numpy arrays of any dimensionality (just like "built-in" numpy functions do). Oh, and if you use set target='gpu', your function also works on GPUs, too. Or you can use target='parallel' to make it parallelize automatically across CPU cores.
- gnufx 8y agoWhen Python was announced, the three main things I remember striking dynamic languages people (apart from using a bastardized offside rule) were weird scope rules, lack of proper GC, and that it appeared to be designed particularly to preclude efficient implementation. We've subsequently seen the huge amount of effort that's been devoted to different ways of working around the implementation issue.
- targafarian 8y agoClaims about Numba by the author seem a little unkind if not wrong to me. Numba can handle vectorized (numpy behavior) directly in addition to explicit loops. The former is accelerated less in comparison to plain python calling numpy (since if you can use numpy operations directly, it's already really fast) but the numpy bits in Numba can also be automatically parallelized by Numba. Explicit loops in numba are accelerated hundreds of times over Python loops (and you can use e.g. prange to write parallel loops, too). Point is, the two paradigms can be mixed and matched at will within Numba. It seems like every example is of cython, but then the author generalizes the conclusions to Numba as well. It would be much more "honest" to show side-by-side comparison of Numba, Cython, and Pythran, since these all have different syntaxes and are fairly different tools. Another example is that you don't have to rewrite functions for different argument types in Numba, but you do in Cython (see "convolve_laplacian" example, which can work with a simple decorator as a numba function). There again, the impression is given that Numba suffers from the same issue as Cython (and as mentioned elsewhere in the comments here, it's possible that Cython has a way around this, but I don't know the details).