13 ms·
Attempts to make Python fast
- ufo 6y agoIIRC Psyco was a precursor to PyPy. Armin Rigo was involved in both.
- beervirus 6y agoPsyco was great. Add two lines of code, suddenly everything is (at least a little, often a lot) faster.
- xioxox 6y agoIt was great. It showed that it is actually possible to run fast Python from within the standard interpreter with excellent compatibility. The only downside I remember was the memory usage.
- chrisseaton 6y agoWhy isn't everyone using it then?
- beervirus 6y agoDoesn’t support modern Python versions. >12 March 2012 >Psyco is unmaintained and dead. Please look at PyPy for the state-of-the-art in JIT compilers for Python. http://psyco.sourceforge.net/ http://psyco.sourceforge.net/
- Naac 6y agoThis article appear to be a list of python interpreters. Not all of these were designed for speed,l. For example jython was also intended for Java/python interoperability. Some of the interpreters on the list haven't seen updates in a while, or don't support python 3.x
- nknealk 6y agoNumba is actually a pretty interesting project. It allows you to JIT compile a single function with a decorator. Static typing required, and it plays nice with numpy. They’ve also got some interesting stuff going on that lets you interface with nvidia GPUs as well. Highly recommend it for anyone doing scientific computing
- sitkack 6y agoI agree, Numba is awesome for lots of reasons. The biggest advantage for everyone, the Numba team as well as its users, is that it is opt-in and done with intent. The programmer is saying, "I am willing to constrain my code to get perf". And that you can do that inside an existing runtime is pretty damn cool. I think for Python to get decent speedups the semantics for the code being optimized needs to be highly constrained. Optimizing full in the wild Python code is a huge huge task. Optimizing for operations over constant type arrays is much much easier. Yes this doesn't speed up the call or the allocation rate, but start with some easy stuff or nothing will improve.
- weakfish 6y agoForgive my ignorance, I'm not super knowledgable on the subject but does this mean you just add decorators to existing functions with typing and it enhances the speed?
- michelpp 6y ago
- 1wd 6y agoOne more: https://github.com/microsoft/Pyjion https://github.com/microsoft/Pyjion
- forgotpwd16 6y agoOn the repo there's also a comparison[1] with some of the other implementations. [1]: https://github.com/microsoft/Pyjion#how-do-this-compare-to- https://github.com/microsoft/Pyjion#how-do-this-compare-to-
- sitkack 6y agoI find it really interesting, that not only did they do the work of creating a JIT using the CoreCLR for CPython, they created a JIT API so that their system is augmenting CPython and not taking it over. Solid engineering. This also means that one could implement an alternative JIT using Rust or OCaml. https://github.com/microsoft/Pyjion#what-are-the-goals-of-this-project https://github.com/microsoft/Pyjion#what-are-the-goals-of-th...
- deleted 6y ago[deleted]
- the_mitsuhiko 6y agoPython is fundamentally not designed to be faster because it leaks a lot of stuff that’s inherently slow that real world code depends on. That’s mutable interpreter frames, global interpreter locks, shared global state, type slots, the C ABI. The only way to speed it up would be to change the language.
- Liquid_Fire 6y agoYou could say the same about JavaScript, but with very heavy investment there are now several implementations that have improved its performance significantly. Also see PyPy, which manages to squeeze a lot more performance out of Python for many use cases without changing the language.
- awestroke 6y agoJS does not have mutable interpreter frames, global interpreter locks, shared global state, type slots, the C ABI.
- dfox 6y agoJS and Python has essentially same data model with everything being at least conceptually built out of dicts. And well, most JS implementations do not have GIL because they are not multithreaded at all.
- baq 6y agotrue and doesn't matter in the context. you can't change (or inspect) the stack frame as an object. you can kind of look at it with Error().stack. this allows the JS JIT to make assumptions that a python compiler simply cannot.
- dfox 6y agoMost of the things that real application code (ie. not something like debugger) can accomplish by modifying or event inspecting frame objects are going to depend on various things that are documented as CPython implementation details. Also JS runtime that supports some kind of debugger interface has to solve the same class of problems. And the most straightforward solution is not even that complex: you simply have to track when some kind of assumption gets broken and then fall back to interpretation or recompile the relevant code (the most complex part of that probably is converting the native stack frame back into interpreter stack frame, which you have to be able to do anyway in order to even expose it to user code for it to be able to modify it). In fact I think that there are many relatively simple modifications that would make CPython significantly faster, but many such things conflict with each other in ways that make the resulting complexity not worth it.
- willseth 6y agoThe list should probably also include mypyc: https://github.com/python/mypy/tree/master/mypyc https://github.com/python/mypy/tree/master/mypyc
- bastawhiz 6y agoUltimately, at least IMO, no attempt to speed up python will succeed until the issue of Python's C API is addressed. This is arguably Pypy's only major barrier: if you can't run the software on it, you're not going to use it. Pyston was arguably the most serious attempt at fast python while maintaining compatibility with the API, but DBX clearly didn't see the RoI they were hoping to. It's looking like HPy is going to (hopefully) solve this. But finishing HPy and getting it adopted is likely to be a pretty massive undertaking.
- travisoliphant 6y agoI think this is true. I have used the Python C-API heavily having started SciPy and NumPy and Numba. I have a pre-alpha plan for addressing the C-API by introducing EPython (a typed subset of Python for extending it). It is not usable and in idea stage only, but I welcome collaborators and funders: https://github.com/epython-dev/epython https://github.com/epython-dev/epython. Here is a talk that describes a bit more the vision: https://morioh.com/p/6db365736476 https://morioh.com/p/6db365736476
- seg_lol 6y agoInteresting, I assume you are familiar with Terra, Titan and Pallene research languages? I love the idea of typed base language to implement a higher level more flexible language while still being able to drop down for correctness and speed. Gradually dynamically typed, ;) Another thing to look at is https://chocopy.org/ https://chocopy.org/ a typed subset of Python for teaching compilers courses. Might be worthwhile pinging Chocopy students and enticing them towards epython. What is the semantic union and intersection between EPython and Chocopy? [1] http://terralang.org/ http://terralang.org/ [2] https://github.com/titan-lang/titan https://github.com/titan-lang/titan [3] https://github.com/pallene-lang/pallene https://github.com/pallene-lang/pallene
- Rotareti 6y agoThis looks interesting! I think the approach where a typed subset of Python is used to compile a fast extension module is the way forward for Python. This would leave us with a slow but dynamic high-level-variant (CPython) and typed lower-level-variant (EPython, mypyc & co) to compile performant extension modules, which you can easily import into your CPython code. The most prominent of such projects I know of is mypyc [0], which is already used to improve performance for mypy itself and the black [1] code formatter. I think it would be interesting to see how EPython compares to mypyc. [0] https://github.com/python/mypy/tree/master/mypyc https://github.com/python/mypy/tree/master/mypyc [1] https://github.com/psf/black/pull/1009 https://github.com/psf/black/pull/1009
- joncatanio 6y agoNot trying to self-promote, but this might be of interest to you. It's not a fully flushed out implementation, but my project analyzed specific language features that affect performance: https://github.com/joncatanio/cannoli https://github.com/joncatanio/cannoli
- ramraj07 6y agoCan't stop laughing at the most germane name that project could ever have.
- joncatanio 6y agoNeeded something to chuckle at during my work ha!
- hydroxideOH- 6y ago> Leave the features: Take the cannoli Now that's how you title a thesis paper.
- Twirrim 6y agoAnother one missing from that list is Graalpython, https://github.com/graalvm/graalpython https://github.com/graalvm/graalpython. It's in early stages of implementation, aimed at being python3 on top of GraalVM.
- sethgecko 6y agoYuri Selivanov tweeted yesterday that Python 3.10 will be "up to 10% faster" https://twitter.com/1st1/status/1318558048265404420 https://twitter.com/1st1/status/1318558048265404420
- centimeter 6y agoThat seems pretty small compared to the huge gap between python and basically any compiled language.
- stuaxo 6y agoThat's pretty good, in optimisation, 5% at a time is a good win.
- saeranv 6y agoAm I correct that 3.10 comes after 3.9? How does that make sense, shouldn't it increase to 4.x? Is there an actual 3.1 (coming after 3.0) that this conflicts with?
- theandrewbailey 6y ago10 comes after 9, so 3.10 comes after 3.9. There's no major changes that would warrant 3.x to 4.0. It's just the 10th big release after 3.0. Yes, there was a Python 3.1: https://www.python.org/download/releases/3.1/ https://www.python.org/download/releases/3.1/
- eznzt 6y agoVersion numbers are not decimal numbers, they are read like the chapters of a book: 3.10 (chapter 3 section 10) comes after 3.9 (chapter 3 section 9)
- yxhuvud 6y agoWait, python doesn't have any method lookup caching before this? I would have expected that developers looked at what other similar languages are doing, but apparently not enough.
- intrepidhero 6y agoWhat I really want for python is a knob to improve startup time. I've imagined there must be a way to "statically link dependencies so that import isn't searching the disk but just loading from a fixed location/file. There doesn't seem to be many resources on the net. I've found this one: https://pythondev.readthedocs.io/startup_time.html https://pythondev.readthedocs.io/startup_time.html. I tried using virtualenvs to limit my searchable import paths, and messed around with cython in effort to come up with a static linked binary. But I've yet to come up with anything that really improves the startup time. Clearly I have no idea what I'm doing.
- korijn 6y agoHave you established that searching for modules is slow? I think it just takes time to actually process the imported modules and load everything into memory.
- formerly_proven 6y agoOn Windows (what with atrocious NTFS performance and all) an interpreter that's using a zipped library is way faster than one using loose modules.
- intrepidhero 6y agoNo. In fact my experiments suggest otherwise. That was just where my intuition lead me.
- jakear 6y agoI once got quite a bit of startup time improvement by simply swapping out cpython's malloc calls for a version that took a large amount of resources at first (~5GB, can be tuned to your workload), and allocated from that. CPython makes many many thousands of mallocs at startup so this gave significant improvement.
- kirubakaran 6y agoCan you please share some numbers if you have them? How much improvement, etc.
- overgard 6y agoI don't know much about the other ones, but I think you'd have to say PyPy has been a success. Although to be honest, I don't know why it would be better to modify CPython vs. just using PyPy -- the JIT speedup does come with some tradeoffs (memory usage, warmup times), so it seems better just to leave that decision up to the user?
- Anka33 6y agoSpace is not a legitimate block delimiter.
- est 6y agoThe HotPy listed by OP is done by Mark Shannon, the same person of today's proposed 5x speedup Also, some relevant old post: https://news.ycombinator.com/item?id=17107047 https://news.ycombinator.com/item?id=17107047
- zellyn 6y agoForgot one : https://github.com/google/grumpy https://github.com/google/grumpy
- arc776 6y agoI gave up trying to make Python fast since to do so you give up what makes Python good and end up writing C/Cython. On top of this, distributing Python is just... gross, at least for my use cases. Eventually I found Nim and never looked back. Python is simply not built for speed but for productivity. Nim is built for both from the start. It's certainly lacking the ecosystem of Python, but for my use cases that doesn't matter.
- acomjean 6y agoI tend to use Python for batch jobs and things where its speed isn't that important to me. Am I alone in this? When I reach for python its not for speed. Its because its fairly easy to write and has some good libraries. Either its done in a few seconds, or I can wait a few hours as it runs as a background slurm task.. I feel like there is a group that wants python to be the ideal language for all things, maybe because I'm not in love with the syntax, but I'm ok having multiple languages.
- nemothekid 6y agoMany people don't start with Python for speed. They are exactly like you - they write a script that is done in few seconds. Then the data scales, then it takes a few minutes. Then you need it to be faster, and now you either need to rewrite the script. It would be helpful if you didn't need to make this choice.
- deleted 6y ago[deleted]
- Boxxed 6y agoWhatever happened to psyco? I remember it pretty much just working without any hassle and actually providing a noticeable speedup. All the mindshare is now on PyPy -- it's received enormous amounts of engineering and still seems very rough around the edges.
- deleted 6y ago[deleted]
- thelazydogsback 6y agopsyco worked well for me at the time as well -- I remember doing something with it and pyGame, FWIR.
- thelazydogsback 6y agoIt amazes me that the stack-entwined implementation with the GIL remained the canonical version this whole time -- I would think that the Stackless version (or similar) would have been the default long-ago. This really should have made it worth it from a 2.x to 3.x version perspective, even if many people had to rewrite their extensions, and even if some monkey-patching were removed from the language in favor of more disciplined meta-programming.
- beagle3 6y agoStackless still uses the GIL; But it avoids using the C stack most of the time, which opens the door to green threads (of which you can have a lot more than OS threads), suspending processes (dump/undump style, except portably), coroutines and more. There were a couple of GIL-less variations, but they were either incredibly slow, or suffered serious compatibility problems (and often both).
- zanellia 6y agoIn my opinion there is some potential there. Especially exploiting the increasing integration of typing-oriented features (i.e. type annotations) and the interest in using those to carry out static analysis (e.g. in mypy, but also Facebook's Pyre and Microsoft's Pyright and many other), it might be possible to speed up execution times a bit. This is especially true if we restrict the attention to a restricted subset of Python as, e.g., within domain specific languages. It might not make sense to entirely reverse engineer a language that was designed to be duck-typed into a statically typed one. However, for some domain specific applications I find performance oriented static analysis an interesting tool. To make it more concrete, here is an experimental DSL for embedded high-performance computing that uses static analysis and source-to-source (Python-to-C, actually) code transformation: https://github.com/zanellia/prometeo https://github.com/zanellia/prometeo.