13 ms·
Why is Python slow
- alt_ 10y agoDatabase died. Google cache: http://webcache.googleusercontent.com/search?q=cache:5V0TMa0cMecJ:blog.kevmod.com/2016/07/why-is-python-slow/+&cd=1&hl=en&ct=clnk&gl=uk http://webcache.googleusercontent.com/search?q=cache:5V0TMa0... The gist of it is: * Python spends almost all of its time in the C runtime This means that it doesn't really matter how quickly you execute the "Python" part of Python. Another way of saying this is that Python opcodes are very complex, and the cost of executing them dwarfs the cost of dispatching them. Another analogy I give is that executing Python is more similar to rendering HTML than it is to executing JS -- it's more of a description of what the runtime should do rather than an explicit step-by-step account of how to do it. Pyston's performance improvements come from speeding up the C code, not the Python code. When people say "why doesn't Pyston use [insert favorite JIT technique here]", my question is whether that technique would help speed up C code. I think this is the most fundamental misconception about Python performance: we spend our energy trying to JIT C code, not Python code. This is also why I am not very interested in running Python on pre-existing VMs, since that will only exacerbate the problem in order to fix something that isn't really broken.
- lucb1e 10y agoSorry for off topic, but I just tried getting the Google cache as well and it just doesn't work. Both in Firefox and Chromium, when I type "cache:<Ctrl+V><Enter>" in the google.com search bar, nothing happens. I checked the developer console, no network requests are made. It says in grey below the search box "press enter to search", but neither enter nor clicking the blue search icon does anything whatsoever. In the past one could manually type /search?q=cache:something in the address bar and it would force the search, but these days it doesn't seem to work anymore. The search request is done in javascript and the /search?q=x link just prefills the search box, leaving it to javascript to fire the actual query (which then fails to do so). Edit: found a way: disable Javascript. This forces Google to search immediately. No plugin necessary: in the Firefox developer console, use the cog wheel on the right top, then in the right column somewhere near the center you can tick "Disable Javascript". Loading the Google home- and search page is also noticeably faster, by the way.
- premium-concern 10y agoI'm not sure your assumptions support your conclusions here, especially your believe that a JIT compiler wouldn't make inroads on the slow C code being executed. The biggest problem of Python is that it lacks the experts that could write those fast runtimes, and it fails to attract them after the Python leaders declared the GIL to be a non-issue.
- gshulegaard 10y agoThe BDFL (Benevolent Dictator For Life) Guido Von Rossum himself put forth the idea that he would consider a patch removing the GIL in 2007 [1]: > "... I'd welcome a set of patches into Py3k only if the performance for a single-threaded program (and for a multi-threaded but I/O-bound program) does not decrease." Unfortunately, experiments thus far have not succeeded to meet these requirements. There is some work being done by the Gilectomy project to try and meet this bar as well as some other requirements currently though [2]. But it is currently grappling with the afore-discovered performance issues that come with removing the GIL. Also at PyCon 2016, Guido himself mentions the Gilectomy project and it's potential consideration (if it works) for Python 3.6+ [3]. So when you say Python leaders declared the GIL a "non-issue", I think you are oversimplifying the actual reality of what removing the GIL means and why leaders (like Guido) have been reluctant to invest resources pursuing. [1] http://www.artima.com/weblogs/viewpost.jsp?thread=214235 http://www.artima.com/weblogs/viewpost.jsp?thread=214235 [2] https://www.youtube.com/watch?v=P3AyI_u66Bw https://www.youtube.com/watch?v=P3AyI_u66Bw [3] https://youtu.be/YgtL4S7Hrwo?t=10m59s https://youtu.be/YgtL4S7Hrwo?t=10m59s
- brianwawok 10y agoSo Python runs in very low memory because of no JIT. You can run say 10 copies of your Python Web server in the space of 1 Java runtime. Will the Java runtime still win? Sure the benchmarks prove it. But there are advantages to running in low memory. Really cheap hosting for low traffic stuff comes to mind... java on a cheap host tends to be disaster.
- pas 10y agoNot because of JIT, but because of refcounting GC. Also the JVM probably loads (and does, or at least is prepared to do - http://www.azulsystems.com/blog/wp-content/uploads/2011/03/2011_WhatDoesJVMDo.pdf http://www.azulsystems.com/blog/wp-content/uploads/2011/03/2... ) too much crap. (And probably the jigsaw and modular and whatever OpenJDK projects will help with that.) Furthermore, Python is not really memory prudent either. PHP is much better in that regard (very fast startup time, fast page render times, but a rather different approach).
- Blaisorblade0 10y agoAre you the OP author, or working on Pyston? I have basically two questions/curiosities — I'm not asking adversarially: 1) For which code is the C runtime most expensive? Typical Python code tries to leave heavy-lifting in libraries, but what if you write your inner loop in Python? Enabling that is (arguably) one goal of JIT compilation, so that you don't need to write code in C. 2) What about using Python ports of performance-sensitive libraries? In more detail: I arrived at https://lwn.net/Articles/691243/ https://lwn.net/Articles/691243/, but I'm not sure I'm convinced. Or rather: with a JIT compiler you probably want to rewrite (parts of) C runtime code into Python so you can JIT it with the rest (PyPy has already replaced C code in their implementation, so maybe there's work to reuse). For instance, an optimizing compiler should ideally remove abstractions from here: import itertools sum(itertools.repeat(1.0, 100000000)) Optimizing that code is not so easy, especially if that involves inlining C code (I wouldn't try, if possible), but an easier step is to optimize the same code written as a plain while loop. Does Pyston achieve that? I guess the question applies to the LLVM-based tier, not otherwise. Yes, Python semantics allow for lots of introspection, and that's expensive — but so did Smalltalk to a large extent. Yet people managed, for instance, to not allocate stack frames on the heap unless needed (I'm pointing vaguely in the direction of JIT compilers for Smalltalk and Self, though by now I forgot those details).
- masklinn 10y ago> Optimizing that code is not so easy, especially if that involves inlining C code (I wouldn't try, if possible) Inlining through C code probably isn't really an option[0], but the optimisation itself shouldn't be that much of an issue, the rust version and the equivalent imperative loop compile to the exact same code: https://godbolt.org/g/OJHIwc[1] https://godbolt.org/g/OJHIwc[1] [0] unless you interpret — and can JIT — the C code with the same underlying machinery as Truffle does [1] used iter_arith for sum(), but you can replace sum() by an explicit fold for no difference: https://godbolt.org/g/R1BgQQ https://godbolt.org/g/R1BgQQ
- gsnedders 10y ago> [0] unless you interpret — and can JIT — the C code with the same underlying machinery as Truffle does Pyston can easily JIT the C code because it uses LLVM for it's main JIT tier.
- smegel 10y ago> This means that it doesn't really matter how quickly you execute the "Python" part of Python. Another way of saying this is that Python opcodes are very complex, and the cost of executing them dwarfs the cost of dispatching them. That doesn't really explain why Python is slow. Your just explaining how Python works. Why should C code be slow? Usually it is fast. Just saying the opcodes are complex doesn't really help, because if a complex opcode takes a long time, it is usually because it is doing a great deal. Java used to have the opposite problem. It was doing too much at the "Java bytecode" level, such as string manipulation - so they added more "complex" opcodes written in C/C++ to speed things up, significantly. What you really need to explain is why Python is inefficient. Bloated data structures and pointer hopping for simple things like adding two numbers may be a big reason. I know Perl had many efficiencies built in, and was considered quite fast at some point (90s?).
- englishgrammar 10y ago> Your just explaining You meant, you are just explaining. Or you're just explaining. Your means ownership, your cat, your computer.
- mlvljr 10y ago>> pointer hopping Betting on this (and the optimizer being really simple-minded -- at least given how people can make it produce obviously inefficent code in SO examples)
- kamaal 10y ago>>I know Perl had many efficiencies built in, and was considered quite fast at some point (90s?). There are a lot of threads in Perlmonks that talk in detail about speeding up Perl, related project et al. To be summarizing it. Languages like Perl and Python are slow because they do a lot of work out of the box that languages like C don't. There fore when you talk of talk of translating Python to C, or Perl to C. Essentially what you are talking of is translating all that extra action back into C, which will run as fast as Perl or Python itself. The more you make it easy for the compiler to interpret the faster it can run and vice versa. Python is slow for the very reason its famous, its easy for the programmer.
- StavrosK 10y agoBetter-formatted link: https://www.pastery.net/chjxpt/ https://www.pastery.net/chjxpt/
- the8472 10y ago> When people say "why doesn't Pyston use [insert favorite JIT technique here]", my question is whether that technique would help speed up C code. That's a silly question. JITs have knowledge instructions relate to each other. The "C code" you're talking about is not opaque to them, it has meanings that can be optimized in relation to other instructions. When you have a "* 2" bytecode + argument it doesn't just dispatch to a C function that multiplies the input by 2. The compiler knows the semantics of that and can convert it to a shift if appropriate. It's not a JIT's responsibility to "speed up the C code". The C code is part of the interpreter, JITs (generally) don't invoke interpreter functions. Or put differently, if you're spending too much time in the C runtime then maybe more code needs to be ported into JITed language itself so that code can also benefit from JIT optimizations.
- hacknat 10y agoAlso, Python has a global interpreter lock so it has no parallelism.
- gshulegaard 10y agoNo true multi-threading, but multi-processing is not affected by the GIL. Also, outside of parallelism, Python has good concurrency support with Gevent and now AsyncIO in Python 3.
- bluejekyll 10y agoNot none, poor...
- chrisseaton 10y agoI don't think the GIL is ever released while running Python code is it? So there is no parallelism between Python threads.
- poooogles 10y agoYou're forgetting multiprocessing. That manages to get round this by running multiple Python interpreters. Big problem is passing objects is dog slow as everything has to be pickled either way.
- tonyarkles 10y agoSome experiments I did during my coursework for my M.Sc. indicated that it was actually worse-than-useless at the time (2009?) Here's what would happen: - When you've got one CPU core, the threads basically just act like a multiplexer. One thread runs for a while, releases the GIL, and the next thread runs for a while. Not a big deal. - When you've got multiple CPU cores, you've got a thundering herd. When the lock is released, the threads waiting on all of the other cores all try to acquire the lock at the same time. Then after one thread has run on, say, core 3, it's gone and invalidated the cache on the other cores (mark & sweep hurts caches pretty badly). The thundering herd stampedes again and the process continues. - To make matters even worse, each core runs at low utilization (e.g. a quad core machine, each core runs at ~25%). If you've got CPU throttling turned on (which my laptop, where I started the experiments, did), then the system detects that the CPU load is low and scales down the clock speed. Normally, this would result in increased CPU utilization, which would speed the CPUs back up again. Unfortunately, the per-core utilization stays pegged at 25% and things never speed back up again. The system looks at it and says "huh! only 25%! I guess we've got the CPU speed set properly!" Maybe it's gotten better since then? I haven't checked recently. Edit: I wish I had the results handy. The basic conclusion was that you got something like a 1.5x slowdown per additional CPU core. That's not how it's supposed to work! Using taskset to limit a multi-threaded Python process to a single core resulted in significant speedups in the use cases I tried.
- radarsat1 10y agoThat's a silly explanation. A smart JIT would be able to take successive Python opcodes and optimise them together as one operation, so how much time it spends interpreting individual instructions has nothing to do with the "speed", in the sense of "current speed" vs "potential speed". If you want to talk about speed in any meaningful sense, you have to talk about potential for optimisation. It's possible that Python (and its opcodes) are designed in such a way that there is little potential for optimisation. This has to do with the semantics of the language, not how much time it spends in one part of the code or the other. I don't know ultimately how much potential for optimisation Python has, but clearly it's a very difficult problem, so we can say with some certainty that Python is "slow", in the concrete sense that there are no low-hanging fruit left for speeding it up. Edit: I say this as an avid Python user by the way. Especially combined with ctypes, I find the Python interpreter to be an absolutely excellent way to "organize" a bunch of faster code written in C/C++. I actually don't have any problem with Python itself being slow, I kind of like it that way personally. It's easy to understand its execution model if the interpreter is kept fairly simple, and this makes it easy to reason with. But then again I am not writing web-scale backends with it, I am just, more or less, using it to batch calls to scientific C functions. So it really depends on your use case. While I've spent plenty of time tuning every little performance gain out of a tight computational loop in C, I can't think of a single time where I've struggled to figure out how to speed up my Python code -- I just am not using it in ways that that would be necessary.
- ankrgyl 10y ago> That's a silly explanation. A smart JIT would be able to take successive Python opcodes and optimise them together as one operation, so how much time it spends interpreting individual instructions has nothing to do with the "speed", in the sense of "current speed" vs "potential speed". I think you might actually be agreeing with the explanation. The point about big opcodes means that the opportunities to look at a sequence of opcodes (i.e. the Python part) are reduced because you're doing a lot of computation over a relatively small number of opcodes. So the challenge involves optimizing the "guts" of the opcodes and the sequence of "guys" across relatively few opcodes. Their approach to solving this is discussed in the original blog post (https://blog.pyston.org/2016/06/30/baseline-jit-and-inline-caches/ https://blog.pyston.org/2016/06/30/baseline-jit-and-inline-c...). This complication happens to make optimizing Python via JIT compilation a tough problem.
- deleted 10y ago[deleted]
- kaushiks 10y agoI think this post misses the point. Having to enter the runtime in the first place is the problem (and making runtime code marginally faster is not the solution). Fast VMs for dynamically typed languages (Chakra, V8) design the object model and the runtime around being able to spend as much time in JITted code as possible - the "fast path".
- bellajbadr 10y agoError establishing a database connection
- harryf 10y agoWhy is database slow
- elcapitan 10y agoWhy is website architecture slow
- ksec 10y agoTo this day we still have wordpress blog without caching plugin ?
- deleted 10y ago[deleted]
- 0xmohit 10y agoAt least it doesn't emit the entire stack trace along with the message! Question: are static websites so hard to do?
- kowdermeister 10y agoNot everybody is comfortable updating, managing it if you refer to Jeckyll or similar generators.
- anc84 10y agoStatic websites take away one of the crucial pieces of blogging: Comments.
- liw 10y agohttp://ikiwiki.info/ http://ikiwiki.info/ supports comments. There are several variants of "static website": * Site is generated by its author, and won't change until author changes and re-generates it. * Site generates content when something changes (e.g., new comment), not on each page view. (Cf. ikiwiki in cgi mode.)
- deleted 10y ago[deleted]
- jacquesm 10y agoPython is high level glue between efficient C functions. It's blindingly fast if you are allowing those efficient C functions to do the heavy lifting. That's why you can do image processing in python, right up until the point where something you need to do doesn't have a ready made primitive and then your program will slow down tremendously if you don't take the time to write the thing you're missing in C.
- sitkack 10y agoAs glue, it isn't particularly efficient, both in runtime or in programmer affordances. Ctypes and cffi are both fairly new and not very friendly. The number of programmers who "drop to native" is ridiculously small. A better glue language would make this almost transparent.
- vegabook 10y agoevery python programmer who has every used Numpy or Pandas, and in my opinion this is where Python shines and why it's so huge in scientific programming, is "dropping into native". So actually a large amount of people are doing so. And I find Python to be an excellent glue language with almost anything I can think of being possible, and much of my heavy lifting being extremely efficient, especially if you use a modern AVX-enabled Numpy. Arguably anybody who ever accessed a database in Python is also "dropping into native". That's why no sane database is written in Python, but plenty of database-using applications are. The only language I have discovered that approaches the efficiency of Python as "glue" is R, but it's about 20x slower, and doesn't even try to be threaded (which can be a big problem for IO sensitive glue tasks).
- sitkack 10y agoBecause extensions that use native code exist doesn't mean that low friction affordances exist for the median programmer to use native code in their applications. I think we will start to see some interesting projects in this space after Python 3.6 ships.
- dr_zoidberg 10y ago
- grx 10y agoWhy is website down
- max_ 10y agoToo much traffic from HN, the database(probably SQL) could not make as many concurrent connections. Hence the scaling problem.
- deleted 10y ago[deleted]
- kensai 10y agoIt has probably been "slashdotted" (in our time: hackernewsed). :D https://en.wikipedia.org/wiki/Slashdot_effect https://en.wikipedia.org/wiki/Slashdot_effect
- kmod 10y agoSomething in my server's configuration eats more and more memory, until the OOM killer decides that killing the MySQL database looks like a dandy way to reclaim memory.
- poooogles 10y agoTry Varnish, for stuff like this it's pretty perfect.
- saboot 10y agoI never quite grasped the actual machinations of Python until I watched Philip Guo's lectures on Python internals. https://www.youtube.com/watch?v=LhadeL7_EIU https://www.youtube.com/watch?v=LhadeL7_EIU It's a bit long, and definitely over several sittings, but I feel like I really understand Python better and relevant to the post, the complexities and tracking (frames, exceptions, objects, types, stacks, references) that occur behind the curtain which drive Python's slow native performance.
- amelius 10y agoIs the slowness due to the structure of the language, or is it because of the implementation? I guess the latter, because PyPy performs a lot better I hear.
- jerf 10y agoPyPy performs better, but when you perform 2-3x times better than something ~40x slower than C, you still don't end up with a "fast implementation". Just, "not as slow". If you've got Python code in hand and you want it to go faster, PyPy can have a great bang-for-the-buck, but if you want it to be legitimately approaching the limits of the capabilities of the hardware, you'll need a different approach. But let me once again underline that if you have Python code in hand, and you want it to be faster, PyPy is a great option. I'm not being critical of PyPy. A common mantra I've heard dozens of times in the last ~20 years is that there's no such thing as a slow language, only slow implementations. But after witnessing the effort to create "fast" implementations for a lot of slow languages over the past 10 years, and seeing so many of them plateau out at about 10x slower than C, I no longer believe this. Or at least, I no longer believe it is practically true. If there is an implementation of Python somewhere in theoretical program space that is as fast as C, it does not appear to me that it will be possible for humans to produce it.
- bluejekyll 10y agoI agree. Though speed of the JVM for instance is not quite as bad. C comes at a development cost, Rust makes this better, but the memory management is still something that you have to get comfortable with. The question that really nags at me is why do people want interpreted languages in all of these cases? When you're deploying code, you inevitably go through a series of steps in deployment where throwing in a compile wouldn't destroy the workflow. I think for many of these cases, the GIL is a great example of this, the language has over-optimized for development at the cost of its runtime.
- 0xmohit 10y agoThe following may also be of interest: - Why Python is Slow: Looking Under the Hood [0] - Fast Python, Slow Python by Alex Gaynor [1] [0] https://jakevdp.github.io/blog/2014/05/09/why-python-is-slow/ https://jakevdp.github.io/blog/2014/05/09/why-python-is-slow... [1] https://www.youtube.com/watch?v=7eeEf_rAJds https://www.youtube.com/watch?v=7eeEf_rAJds
- SFJulie 10y agoWell, C/Fortran/C++/ASM are a PITA when it comes to dynamic structures like the one for handling configurations. Plus authors missed that boxing in python has a tendency to fragment data in a non controlable way in the memory thus making the use of L1/L2/L3/mem (on x86) or other memory architecture with similar layout very hard to use. If you code often you know the 80/20 rule: 80% of your code is the setup preparing for the 20% of heavy lifting. Numpy (which relies on Fortran) is a nice Proof that when done correctly python is really useful. A computing a Moving average in pure python is 10 000 times slower than using numpy (especially if you use FFT). So ... I am saying since python (like Tcl or Perl) is a good language for doing FFI (foreign function integration) it should be used this way. And thanks to the often unfairly hated GIL it enables to use un-thread safe foreign language library in a thread safe way. All being said and done, if python used this way is slow, I dare say it is because some coders do not understand how to build there data/execution flow. And this is not language dependent, but a question of coder. I thought they were 10x coders in the past. I recently realized there are /10 coders in fact. The one that are pissed at coding when «it does not work the way it should» and expect language to be magically doing most of the job without learning. The 1x coders on the other hands are boring, slow coders and accept that when the «stuff» is not behaving the right way, it may not be the stuff that is not working, but him/her that have a misconception. 1x coder is not a state, it is a trajectory that can degrade or improve with new challenge and poor/good state of mind, hence the misconception on 10x.
- lokedhs 10y agoIt is possible to make a dynamic language natively compiled and very fast. Common Lisp is arguably even more dynamic than Python, and SBCL manages to generate code that rivals C in performance.
- PeCaN 10y agoI don't consider Common Lisp “more dynamic” than Python; at least in that you modify code at runtime and such. Frequently Common Lisp code is not particularly dynamic, because you can get the same level of expressiveness without relying on mucking around with magic at runtime. Also SBCL only generates really fast code if you use a lot of type hints and (optimize (speed 3) (safety 0)). That said, when you do, it's really fast.
- merb 10y agoActually I never found Python "slow". I found it "slow" in some things, but not "all things are slow with python". What was problematic was pulling big lists from a database and doing some stuff with the list. Also it was really akward that "threading" is not really great on python. Especially not when you are using 16 core servers. I mean you could create vm's for that or dockerize that. but that means deployment complexity increases which wasn't our goal. But I've seen a lot of successful python deployments and if you have enough manpower you pretty sure can run with CPython just fine.
- chrisseaton 10y agoWhat you mean is that Python is fast enough for your purposes, which is great for you but it's not what this post is about. I think it's pretty non-controversial to say that basic language operations in Python are slower than in other language implementations, which is what this post is talking about.
- brianwawok 10y agoOn a 16 core server you can run 32 copies of your program Ala Celery. It will peg the CPUs just dandy..
- retrogradeorbit 10y agoMy experience is that it is not "dandy". We have had a lot of trouble pegging the CPUs on our celery worker boxes (doing CPU bound jobs). You get more than one CPU utilised, sure, but we never can seem to get all cores fully utilised. We rewrote some of the tasks into a single multi-threaded JVM process pulling off the rabbit queue and they instantly and consistently pegged every CPU at 100%. I wish I knew how to get our celery worker farm to full utilisation because it would save us a fair bit of money.
- jstanley 10y agoHe's saying run 32 individual processes, not 32 threads within one process. Python's global interpreter lock will knobble you if you're using threads.
- max_ 10y agoWhy can't someone just write another compiled language with 100% Python syntax? e.g Julia. only more general purpose.
- thebooktocome 10y agoJulia is general-purpose already.
- max_ 10y agoI thought they were focused towards technical computing?
- ZenoArrow 10y agoThat sort of focus is expressed by the library ecosystem, not the language itself. You can write any type of program you want with Julia, you'll just be doing more work if you're building something that can't fully rely on preexisting libraries. As the library ecosystem becomes broader and more mature, this issue goes away.
- aidos 10y agoBecause the libraries keep people on Python and you'd need to support all those too. Edit to add, my understanding is that the flexible nature of Python makes it hard to swap out the python without implementing all the stuff that makes it flexible. In a way I guess pypy is the python you're talking about. And it does exist, and is faster, but doesn't support all libraries.
- chrisseaton 10y agoMaybe some people use Python because of the semantics. If you just replicate the syntax would it really be Python any more?
- coldtea 10y agoYou can't separate the (full) syntax from the semantics, so he means both.
- eldude 10y agoCan someone ELI5 why Python is more like rendering HTML than executing JS? This is confusing to me since much of node.js/V8 is C, yet AFAIK (from the title and my experience) it's faster, and I don't recall anything intrinsically more declarative (i.e. HTML) when writing python compared to JS. They both feel very similar as scripting languages to me. IOW, from my limited ignorant perspective, this feels more like the WHAT than the underlying differentiating WHY. It's possible it's in there and I missed it. EDIT: FWICT from the linked slides[1], it's the result of 2 issues: 1. Expensive dynamic language features and 2. python is like node.js, but as if you only called V8 bindings and so VM performance was irrelevant. This is strange to me; while I feel I can conceptualize the difference, I still don't know enough to understand why it is so compared with node.js. [1] https://docs.google.com/presentation/d/10j9T4odf67mBSIJ7IDFAtBMKP-DpxPCKWppzoIfmbfQ/mobilepresent?slide=id.g142ccf5e49_0_453 https://docs.google.com/presentation/d/10j9T4odf67mBSIJ7IDFA...
- dvogel 10y agoThe HTML comparison was referring to the python bytecode rather than the language. A more apt comparison would be CISC vs RISC processors, where the V8 IR is more like a RISC processor.
- chrisseaton 10y agoI think it's true that language implementations such as Ruby and Python spend most of their time running the C parts of the code. I did a talk saying the same thing about Ruby a couple of weeks ago, but referring to the Java code in JRuby, https://ia601503.us.archive.org/32/items/vmss16/seaton.pdf https://ia601503.us.archive.org/32/items/vmss16/seaton.pdf. But this doesn't mean that a JIT is not going to help you. It means that you need a more powerful JIT which can optimise through this C code. That may mean that you need you to rewrite the C in a managed language such as Java or RPython which you can optimise through (which we know works), or maybe we could include the LLVM IR of the C runtime and make that accessible to the JIT at runtime (which is a good idea, but we don't know if it's practical). I work on an implementation of Ruby, and we make available the IR of all our runtime routines (in our case implemented in Java) to a powerful JIT, so that we can inline from the interpreter into the runtime and back again. In the case of Python, PyPy does the same thing, allowing the JIT to optimise between the interpreter and runtime, as they're both written in RPython. So I think the problem the Pyston project needs to solve is how to allow the JIT to see the runtime routines and optimise through them like it does with Python code.
- brianwawok 10y agoPypy makes your app take many times the memory for like 20% perf. Which is good but seems maybe often not worth the effort.
- adrianN 10y agoOn Pypy's benchmark site, the speedup is a lot higher than 20%. I usually experience a 2x-3x speedup with the kind of code I run, sometimes more. > http://speed.pypy.org/ http://speed.pypy.org/
- e12e 10y agoEh... cat<<eof > float.py import itertools s = sum(itertools.repeat(1.0, 100000000)) print(s) $ time python float.py 100000000.0 real 0m0.602s user 0m0.596s sys 0m0.004s time python3 float.py 100000000.0 real 0m0.603s user 0m0.600s sys 0m0.000s $ time pypy float.py 100000000.0 real 0m0.211s user 0m0.088s sys 0m0.004s That's with no warmup for the pypy variant (or indeed the other python variants). Or, slightly more "robust": $ python -m timeit -s "import itertools as i" \ "sum(i.repeat(1.0, 100000000))" 10 loops, best of 3: 594 msec per loop $ python3 -m timeit -s "import itertools as i" \ "sum(i.repeat(1.0, 100000000))" 10 loops, best of 3: 592 msec per loop $ pypy -m timeit -s "import itertools as i" \ "sum(i.repeat(1.0, 100000000))" 10 loops, best of 3: 68.2 msec per loop Pypy actually does pretty good here: $ cat float.cpp #include<iostream> int main() { double s = 0; for (int i = 0; i < 100000000; ++i) { s++; } std::cout << s << std::endl; return 0; } $ g++ --std=c++14 -O3 float.cpp $ time ./float 1e+08 real 0m0.237s user 0m0.236s sys 0m0.000s Note that the C++ code use a loop, not a lazy generator. Apparently they may be coming in c++17 as proposal N4286.
- TempleOS 10y agoGPU vs CPU https://www.youtube.com/watch?v=XcolCeWIcss https://www.youtube.com/watch?v=XcolCeWIcss
- beagle3 10y ago(based on discussion, can't get to website) Any discussion that does not compare to LuaJIT2 is suspect in its conclusions. On the surface, Lua is almost as dynamic as Python, and LuaJIT2 is able to do wonders with it. Part of the problem with Python is (that Lua doesn't share) is that things you use all the time can potentially shift between two loop iterations, e.g. for x in os.listdir(PATH): y = os.path.join(PATH, x) process_file(y) There is no guarantee that "os.path" (and thus "os.path.join") called by one iteration is the same one called by the next iteration - process_file() might have modified it. It used to be common practice to cache useful routines (e.g. start with "os_path_join = os.path.join" before the loop and call "os_path_join" instead of "os.path.join"), thus avoiding the iterative lookup on each iteration. I'm not sure why it isn't anymore - it would also likely help PyPy and Pyston produce better code. This is by no means the only thing that makes Python slower - my point is that if one looks for inherently slow issues that cannot be solved by an implementation, comparison with Lua/LuaJIT2 is a good way to go.
- bakery2k 10y agoPart of the problem with Python is (that Lua doesn't share) is that things you use all the time can potentially shift between two loop iterations Why is this problem not shared by Lua? Does Lua somehow prevent "process_file" from modifying the equivalent of "os.path"?
- xamuel 10y agoThere isn't even any guarantee that "os.path" is the vanilla iterative lookup it's presented as. For all we know, it could be a @property-decorated method thousands of lines long.
- e12e 10y agoFrom the linked [lwn] article: "Another example he gave demonstrates the slowness of the C runtime: import itertools sum(itertools.repeat(1.0, 100000000)) That will calculate the sum of 100 million 1.0s. But it is six times slower than the equivalent JavaScript loop. Float addition is fast, as is sum(), but the result is not. Larry Hastings asked what it was that was slowing everything down. Modzelewski replied that it is the boxing of the numbers, which requires allocations for creating objects. Though an audience member did point out with a chuckle that you can increase the number of iterations and Python will still give the right answer, while JavaScript will not." Reminded me about the excellent talk about Julia: "Julia: to Lisp or not to Lisp?" https://www.youtube.com/watch?v=dK3zRXhrFZY https://www.youtube.com/watch?v=dK3zRXhrFZY One things he points out early is that both the C99 and the R6RS Scheme spec is 20% numerical. Because correct and (reasonably) fast numbers and arithmetic is actually pretty hard to get right on a computer - if you want to abstract away hardware "short-cuts" and allow for precise arithmetic by default. It will be interesting to see how much type hints (eliminating some of the boxing/unboxing) will help python. And if it turns out to really be a good fit for the language -- everyone wants "free" performance, but transitioning to a (even partially) typed language is certainly not "free". Another point with the above loop, is that ideally, even if it can't be optimized/memoized down to a constant - it really shouldn't have to be much slower than its C counterpart. Except for handling bignums in some way or other (perhaps on overflow only). [lwn] https://lwn.net/Articles/691243/ https://lwn.net/Articles/691243/
- hyperpape 10y agoMy understanding is that there is no plan to use type hints for optimization. I'm on a phone so I can't find it now, but I previous talked to some people on here about it in comments.
- e12e 10y agoPlease let us know if you find a reference to that - I've not seen anything either way (other than perhaps eg: cython piggybacking on the syntax for its own use - but not something strictly speaking "in (c)python". That said, according to: https://www.python.org/dev/peps/pep-0484/#rationale-and-goals https://www.python.org/dev/peps/pep-0484/#rationale-and-goal... "This PEP aims to provide a standard syntax for type annotations, opening up Python code to easier static analysis and refactoring, potential runtime type checking, and (perhaps, in some contexts) code generation utilizing type information. Of these goals, static analysis is the most important. This includes support for off-line type checkers such as mypy, as well as providing a standard notation that can be used by IDEs for code completion and refactoring." So optimization doesn't appear to be a non-goal - as much as not a primary goal?
- __s 10y agoYet we got a ~5% speed advantage when I refactored bytecode to wordcode https://bugs.python.org/issue26647 https://bugs.python.org/issue26647
- therestisgone 10y agoVery nice to see some circular reasoning here. Because python is slow everyone dispatches to C code. Because everyone dispatches to C code, optimising the python interpreter is not worth it.
- shadowmint 10y agoPython is made faster by optimizing the C implementation of it's opcodes, not by doing any optimization at a higher level? Really? Sounds more like; It's more convenient for us to try to optimize cpython at the opcode level because its technically difficult to apply JIT techniques at a higher level because of the way cpython is implemented. Pyston's performance improvements come from speeding up the C code, not the Python code. When people say "why doesn't Pyston use [insert favorite JIT technique here]", my question is whether that technique would help speed up C code. I think this is the most fundamental misconception about Python performance: we spend our energy trying to JIT C code, not Python code. This is also why I am not very interested in running Python on pre-existing VMs, since that will only exacerbate the problem in order to fix something that isn't really broken. ...really? All I can fathom is that this project is about trying to take the existing cpython implementation and make it faster by applying various magical hacks at a very low level, rather than trying to address any of the more difficult problems about why the cpython runtime is slow. This is exactly the opposite approach from pypy (ie. reimplement cpython in a way which is fundamentally better); and it certainly seems to be yielding some interesting results. ...but I think I'm a little skeptical that its the only solution. It just happens to be the solution they've decided to pursue.
- known 10y agoAny script language is slow
- pubby 10y agoNot really. LuaJIT and V8 do pretty well on benchmarks. Maybe not "fast", but they certainly aren't slow.
- andybak 10y ago'script language' isn't a technical definition. Do you mean 'interpreted'? 'dynamic'? Even those terms need to be carefully qualified and counter-examples exist in each case.
- rurban 10y agoOnly the old dumb ones (python, perl, ruby, formerly also php). Normal dynamic script languages are written by engineers and are therefore very fast.
- pkd 10y agoAre you implying that people like Larry Wall, Yukihiro Matsumoto and Guido van Rossum, all of who have advanced degrees in computer science are not (good?) engineers?
- PeCaN 10y agoWell, to be fair, the original Ruby interpret was a very naive AST-walking interpreter. Not that Matz isn't a good engineer; I think he probably just wanted to experiment. (Take a look at the Ruby parser sometime if you need confirmation that Matz is both brilliant and possibly insane.) Perl's internals are a little weird but are really fast at text processing. I wouldn't necessarily hold Perl up as a good piece of software engineering, but it's still much better than what most people could write. (I'm agreeing with you though. Your parent's comment is absurd.)
- rurban 10y agoThey might be good hackers, but have no idea how a VM should be implemented. That's why you got this mess. Just look how they designed the ops, the primitives, the calling convention, the stack, the hash tables. This is not engineering, this was amateurish compared to existing practice.
- makecheck 10y agoPython has had some surprising costs but some well-known cases have straightforward work-arounds. For example, at one point, the way you called a multi-dot function actually mattered (not sure if this persists in the latest Python releases). If your function is like "os.path.join", the interpreter would seem to incur a cost for each dot: once to look up "os.path" and once to look up "path.join". This meant that there was a noticeable difference between calling it directly several times, and calling it only through a fully resolved value (e.g. say "path_join = os.path.join" and call it only as "path_join" each time). Another good option is to use the standard library’s pure-C alternatives for common purposes (containers, etc.) if they fit your needs.
- ebbv 10y agoFor me this post isn't an argument against Mozilla implementing a similar dialogue; it's an argument for implementing one and improving it. To really have meaningful improvements would require totally reworking how extensions deal with the DOM and requiring them to ask for different levels of permission. That would break most plugins as written and probably require significant work to implement.
- alephnil 10y agoThe argument is that Python spend most time in the C runtime, but that is really only a property of the current implementation. If I use Jython or even implement Python on top of bare metal, making Python the OS, that will not be spending any time in the C runtime because there wouldn't be one, but that's besides the point. The question is whether an implementation could be made significantly faster than the current Python interpreter is, or whether there are properties of the Python semantics that makes that hard. One such thing is the amount of dynamic behavior that is allowed in Python. That require most variables to be boxed, even for basic types like numbers. There are dynamic languages (LuaJIT was mentioned, but Javascript, Julia and several Lisp implementation could be mentioned as well) that are considerably faster than Python, so why isn't Python fast? Personally I think that Python could be made quite a bit faster than it is, but such a new Python system would almost certainly be incompatible with a large number of the Python libraries that interface with native code. For most Python users, this availability of libraries is a major motivation for using Python in the first place, and a new fast implementation without such compatibility would be worthless for most users.
- dicroce 10y agoimho it's about references (pointers) and the inability of the memory prefetcher to optimize memory accesses. to fix this languages need true value types.
- alayne 10y agoJust because you're in the C runtime, doesn't mean you're doing productive work. If it isn't the bytecode VM that is slow, then why is PyPy able to be much faster in many cases. In my experience, Python is easily 10-20x slower than a compiled language when doing computational work where you can't just call into a big C function, just like you would expect to see from any interpreter. I won't generally use it for anything data intensive.
- rurban 10y agoThe arguments presented here are extremely dumb. Of course everyone knows already that the ops itself are slow. What not many people know is that all this could be easily optimized. javascript had the very same problems, but had good engineers to overcome the dynamic overhead. php7 just restructured their data, lua and luajit are extreme examples of small data and ops winning the cache race. look at v8, which was based on strongtalk. look at guile. look at any lisp. All these problems were already solved in the 80ies, and then again in the 90ies, and then again in the 2000ies. python is similar to perl and ruby plagued with dumb engineering. python is arguably the worst. They should just look at lua or v8. How to slim down the ops to one word, how to unbox primitives (int, bool), how to represent numbers and strings. How to speed up method dispatch, how to inline functions, how to optimize internals, how to call functions and represent lexicals. How to win with optional types. Basically almost everything done in python, perl and ruby is extremely dumb. And it will not change. I'm still wondering how php7 could overcame the very same problems, and I believe it was the external threat from hhvm and some kind of inner revolution, which could convince them to rewrite it. I blame google who had the chance to improve it, after they hired Guido, and went nowhere. You still have the old guys around resisting any improvements. They didn't have an idea how to write a fast vm's in the old days, and they still don't know.
- sitkack 10y agoThe semantics of the language are extremely flexible, that flexibility add a ton of "yeah but" when trying to do obvious things. "yeah but what if they monkey patched boolean?" ... If I were starting a Python VM from scratch, objects would be closed and only deopt if they were mucked with. Most Python code in the wild is written in a dynamically typed version of Java. Yes most Python is Java w/o the types. Typed inference, escape analysis and closed classes will allow for the same speed as the JVM.
- rurban 10y agoYes, but then it's not python anymore. You will get a similar speedup for a fully dynamic python with easier optimizations, as done with v8 or lua. Even without types. types and attributes (like closed classes) do help, but a proper vm design is more important. But not within this community. It needs to be a completely separate project, as cython, mypy or pypy. pypy is pretty good already, but a bit overblown.
- saynsedit 10y agoWhat's expensive about the runtime is the redundant type/method-dispatch not the opcode dispatch. The runtime is constantly checking types and looking up methods in hash tables. Gains can be made by "inter-bytecode" optimization in the same vein as inter-procedural optimization. If you can prove more assumptions about types between the execution of sequential opcodes, you can remove more type checks and reduce the amount of code that must be run and make it faster. E.g.: 01: x = a + b 02: y = x + c If we have already computed that a and b are Python ints, then we can assume that x is a Python int. To execute line 2, we just then need to compute the type of c, thus saving time. The larger the sequence of opcodes you are optimizing over, the better gains you can make. Right now I think Pyston uses a traditional inline-cache method. This only works for repeated executions. The code must eagerly fill the inline cache of each opcode to get the speed I'm talking about. Another reason Python's runtime is slow is because there is no such thing as undefined behavior and programming errors are always checked against. E.g. The sqrt() function always checks if its argument is >= 0, even though it's programming error to use it incorrectly and should never happen. This can't be fixed by a compiler project, it's a problem at the language level. Being LLVM based, I think Pyston has its greatest potential for success as an ahead-of-time mypy compiler (with type information statically available). IMO leave the JITing to PyPy.
- jlarocco 10y agoThe notion that Python spends most of its time in its C code isn't particularly insightful. That's how interpreters traditionally work. It just raises the question of why that C code is slower than other interpreters. Personally, I don't mind Python's speed. I don't use it for runtime speed, I use it for development speed. If I need runtime speed I use C++. Lately, though, I'm getting to the point where I write Common Lisp about as quickly as Python and it typically runs 4-5x faster than Python, so I've just been using that.
- elchief 10y agoWho cares? You're not using it because it's performant. You use it because it's fast to code in and to ship. If you only care about performance, write C or Java or an assembler and extend your ship date.
- collyw 10y agoObviously being down voted for being too pragmatic.
- Animats 10y agoPoor article. This subject has been covered many times, and others have put in the key references. Python's dynamism is part of the problem, of course. Too much time is spent looking up objects in dictionaries. That's well known, and there are optimizations for that. Python's concurrency approach is worse. Any thread can modify any object in any other thread at any time. That prevents a wide range of optimizations. This is a design flaw of Python. Few programs actually go mucking with stuff in other threads; in programs that work, inter-thread communication is quite limited. But the language neither knows that nor provides tools to help with it. This is a legacy from the old C approach to concurrency - concurrency is an OS issue, not a language issue. C++ finally dealt with that a little, but it took decades. Go deals with it a little more, but didn't quite get it right; Go still has race conditions. Rust takes it seriously and finally seems to be getting it right. Python is still at the C level of concurrency understanding. Except, of course, that Python has the Global Interpreter Lock. Attempts to eliminate it result in lots of smaller locks, but not much performance increase. Since the language doesn't know what's shared, lots of locking is needed to prevent the primitives from breaking. It's the combination of those two problems that's the killer. Javascript has almost as much dynamism, but without extreme concurrency, the compiler can look at a block of code and often decide "the type of this can never change".
- EE84M3i 10y agoNote that threading is not the only place this can happen: this can happen as the result of signal handlers as well. cpython has can trigger signal handlers between any python opcodes.
- markhahn 10y agoImportant clarification: when the author says "C runtime", what he means is "Python runtime written in C". The C runtime is, of course, libc, libm, etc. I'm not sure why the author thinks Python's high-level IR (please don't call it a VM) is a good thing, or unoptimizable (to a better IR or native code). Perhaps he's never read about Smalltalk/Self optimization (which is eye-opening!)