15 ms·
Python performance: it’s not just the interpreter
- g8oz 6y agoI'd be interested in seeing how PHP performs running equivalent code.
- kmod 6y agoI don't know PHP, but feel free to send a pull request!
- owyn 6y agoUsing range() in php it takes about 1.3 seconds. The PHP docs imply that it's a generator now but I'm not positive about that. I wrote an equivalent php function just using for loops and calling strval($x) 20 million times and on my laptop it runs in .9 seconds... The second form of creating 20 lists with 1m elements runs in .4 seconds. Without needing to write 300 lines of Py/C stuff. So... shrug? microbenchmarks? I guess writing the optimized code was the fun part and the actual benchmark/timing part doesn't really matter. It's just for loops... the other comments are right that benchmarks definitely matter when you're doing more varied "work" though. For what it's worth, I did benchmark a big application in PHP "for real" and parsing a large configuration file (10,000 lines +) on every request did take up about %15 of the wall clock time. We optimized a few things there because it was worth it. It was a HUGE application of about ~1M LoC and an average request was about 300-400 milliseconds so... I guess it wasn't doing a lot of for loops? Edit: I had a few minutes before my next meeting to code golf this so I decided to test if range() worked with yield [1] <? function gen() { // 1.3s foreach (range(0,20) as $i) { foreach (range(0,1000000) as $j) { strval($j); } } } function foo() { // .9s for ($i = 0; $i < 20; $i++) { for ($j = 0; $j < 1000000; $j++) { strval($j); } } } function bar() { // .4s for ($i = 0; $i < 20; $i ++) { $x = range(0, 1000000); } } // [1] returns in .06 seconds // pretty sure this just executes yield once and does no real work just like me this morning function gen2() { foreach (range(0,20) as $i) { foreach (range(0,1000000) as $j) { yield; } } }
- meritt 6y agoIt's about 3-4x faster with the initial implementation: foreach(range(1,20) as $j) foreach(range(1,1000000) as $i) strval($i); Running on my low-end EC2 box: php7.3 py.php == 1.258s python2 py.py == 3.714s python3 py.py == 4.530s
- carapace 6y agoSite broken, "Hug of Death"? > This page isn’t working > blog.kevmod.com is currently unable to handle this request. > HTTP ERROR 500
- kmod 6y agoThis is what I get for hosting my own wordpress server. It was struggling and I installed a caching plugin which took down the site. Should be back up, sorry!
- carapace 6y agoCheers! (Nice problem to have though, eh?) :) Great blog post.
- SketchySeaBeast 6y agoInstalling plugins into Wordpress kind of feels like doing your own brain surgery.
- anentropic 6y agoWhen it says that "argument passing" was responsible for 31% of the time, do I understand right that we're talking about this line in the inner loop? str(i) ...and the time is spent packing i into a tuple (i,) and then unpacking it again? are keyword args faster? or they do the same but via dict instead of tuple I guess
- kmod 6y agoYep! It's slightly worse than you would think. Here's a (slightly edited) version of how the argument passing works static PyObject * unicode_new(PyTypeObject *type, PyObject *args, PyObject *kwds) { PyObject *x = NULL; static char *kwlist[] = {"object", "encoding", "errors", 0}; char *encoding = NULL; char *errors = NULL; PyArg_ParseTupleAndKeywords(args, kwds, "|Oss:str", kwlist, &x, &encoding, &errors)) } Notice the call to PyArg_ParseTupleAndKeywords -- it takes an arguments tuple and a format string and executes a mini interpreter to parse the arguments from the tuple. It has to be ready to receive arguments as any combination of keywords and positional, but for a given callsite the matching will generally be static.
- anentropic 6y agoAnd is that literally every Python function doing that under the hood? Even a 1-arity function? I don't get what the format string is for. Anyway, thanks for answering!
- kmod 6y agoIt used to be this way, but at some point they added a faster calling convention and moved a number of things to it. Calling type objects still falls back to the old convention though.
- chrisseaton 6y agoKevin knows much more than I do about optimising Python, but aren't lots of the things listed as 'not interpreter overhead' only slow because they're being interpreted? For example you only need integer boxing in this loop because it's running in an interpreter. If it was compiled that would go away. So shouldn't we blame most of these things on 'being interpreted'?
- jsnell 6y agoNot really. Integer boxing is totally orthogonal to compilation vs interpretation. Assuming we need to preserve program semantics, a naive compiler would have exactly the same boxing overhead. And on the other hand a smart interpreter could eliminate it. (E.g stack allocate integers by default, dynamically detect conditions that require promoting them to the heap, annotate the str builtin as escape safe.)
- chrisseaton 6y agoIsn't removing overhead like boxing the main point of a compiler? Seems like if you write a compiler that doesn't do that you might as well not have bothered.
- throwaway894345 6y ago> In this post I hope to show that while the interpreter adds overhead, it is not the dominant factor for even a small microbenchmark. Instead we'll see that dynamic features -- particularly, features inside the runtime -- are to blame. I'm probably uneducated here, but I don't understand the distinction between the runtime and the interpreter for an interpreted language? Isn't the interpreter the same as the runtime? What are the distinct responsibilities of the interpreter and the runtime? Is the interpreter just the C program that runs the loop while the runtime is the libpython stuff (or whatever it's called)?
- jcelerier 6y agoC does not have an interpreter (well, the most used implementations don't) but often has a runtime (libc). It's really a runtime and not just a library, because a C compiler can insert calls to its functions even if they aren't part of your code at all, like here : https://gcc.godbolt.org/z/_w78qd https://gcc.godbolt.org/z/_w78qd
- barrkel 6y agoYou're quite right - the distinction is entirely synthetic here. If you could swap out the runtime without changing the interpreter, the distinction would be far more valid.
- PaulDavisThe1st 6y agoConsider a program written in Python. Imagine it converted into some "internal representation" that will be used by the interpreter to execute it. We can ask questions like "when I write a statement in Python that adds two integers (typically a single machine instruction in a C program), how many machine instructions are executed while running that statement?" Those are questions about the interpreter - essentially, the raw speed and execution patterns of the language. But when you're running a Python program (or one written in many other languages these days), there is a lot going on inside the process that you cannot map back to the Python statements you wrote. The most obvious is memory management, but there are several others. This stuff takes time to execute and can change the behavior of the statements that you did write in subtle (and not so subtle ways). Note that this kind of issue exists even for compiled languages too: when you run a program written in C++, the compiler will have inserted varying amounts of code to manage a variety of things into the executable that do not map back explicitly to the lines you wrote. This is the "C++ runtime" at work, and even though the "C++ language" may run at essentially machine speed, the runtime still adds overhead in some places. Interpreted languages are the same, just worse (by various metrics)
- nojito 6y agoA potential issue with benchmarks like this is that there are instances where the initial findings don't scale. I would be interested to see how it does over an operation that takes 1 minute, 5 minutes, 10 minutes.
- barrkel 6y agoThe article is a fine example of incremental optimization of some Python, replacing constructs that the standard Python interpreter has overheads in executing with others which trigger fewer of those same overheads. The title isn't quite right, though. Boxing, method lookup, etc. come under "interpreter" too. There's a continuum of implementation between a naive interpreter and a full-blown JIT compiler, rather than a binary distinction. All interpreters beyond line-oriented things like shells convert source code into a different format more suitable for execution. When that format is in memory, it can be processed to different degrees to achieve different levels of performance. For example, if looking up "str" is an issue, an interpreter could cache the last function it got for "str" conditional on some invalidation token (e.g. "global_identifier_change_count"), so it doesn't in fact need to look it up every time. Boxing can be eliminated by using type-specific representations of the operations, and choosing different code paths depending on checking the actual type. Hoisting those type checks outside of loops is then a big win. Add inlining, and the hoisting logic can see more loops to move things out of. Inlining also gets rid of your argument passing overhead. None of this requires targeting machine code directly, you can optimize interpreter code, and in fact that's what the back end of retargetable optimizers looks like - intermediate representation is an interpretable format. Of course things get super-complex super-quickly, but that's the trade-off.
- chrisseaton 6y ago> There's a continuum implementations between a naive interpreter and a full-blown JIT compiler And many full-blown JIT compilers also include an interpreter for things like deoptimisation and rolling towards a safe state. Few things are a pure JIT (I think the .NET runtime is) and actually you don't want that is limits how far you can optimise.
- drcongo 6y agoThat was really interesting, even though a lot of it (almost all the C) was way over my head. Thank you.
- mark-r 6y ago> The benchmark is converting numbers to strings, which in Python is remarkably expensive for reasons we'll get into. I was a bit disappointed that converting numbers to strings was the only thing he didn't actually discuss. I've discovered that the conversion function is unnecessarily slow, basically O(n^2) on the number of digits. This despite being based on an algorithm from Knuth.
- pedrovhb 6y ago> And impressively, PyPy [PyPy3 v7.3.1] only takes 0.31s to run the original version of the benchmark. So not only do they get rid of most of the overheads, but they are significantly faster at unicode conversion as well. Wow, that's pretty impressive. I never really got to use PyPy though, as it seems that for most programs either performance doesn't really matter (within a couple of orders of magnitude), or numpy/pandas is used, in which case the optimization in calling C outweighs any others. Can anyone share use cases for PyPy?
- kevin_thibedeau 6y agoAnything where function call overhead is an issue and you don't have a native library escape hatch. Parsers written in Python are inherently slow. Especially so with parser combinators.
- wyldfire 6y ago> Can anyone share use cases for PyPy? If you are concerned about performance, it really shines. It speeds up computational workloads, for sure. But it also improves performance in lots of I/O scenarios too.
- 6c696e7578 6y ago> Can anyone share use cases for PyPy? Well, anything that you need to do where the libraries are there and waiting for you. If you get into the territory of missing libraries it can be a bit of a pain. Otherwise, it's a breath of fresh air as it's almost a drop in replacement.
- hangonhn 6y agoJust to clarify, does it matter if these libraries are pure Python? What I have a hard time understanding is how PyPy is almost a drop-in replacement but then have issues with missing libraries. Couldn't you just pip install the libraries or just literally get the source and run them if they're pure python?
- 6c696e7578 6y ago
- no_gravity 6y agoI wanted to play with variations of the code. For that it is useful to make it output a "summary" so you know the variation you tried is computationally equivalent. For the first benchmark, I added a combined string length calcuclation: def main(): r = 0 for j in range(20): for i in range(1000000): r += len(str(i)) print(r) main() When I execute it: time python3 test.py I get 8.3s execution time. The PHP equivalent: <?php function main() { $r = 0; for ($j=0;$j<20;$j++) for ($i=0;$i<1000000;$i++) $r += strlen($i); print("$r\n"); } main(); When I execute it: time php test.php Finishes in 1.4s here. So about 6x faster. Executing the Python version via PyPy: time pypy test.py Gives me 0.49s. Wow! For better control, I did all runs inside a Docker container. Outside the container, all runs are about 20% faster. Which I also find interesting. Would like to see how the code performs in some more languages like Javascript, Ruby and Java.
- hu3 6y agoThe next PHP release will sport a JIT compiler [1]. I'd wager it will reach similar performance than pypy. [1] https://stitcher.io/blog/new-in-php-8 https://stitcher.io/blog/new-in-php-8
- iruoy 6y agoI've added rust using this code fn main() { let mut r = 0; for _x in 0..20 { for y in 0..1_000_000 { r += y.to_string().len(); } } println!("{}", r); } Surprisingly PyPy is the fastest % hyperfine target/release/perftest "php perftest.php" "python perftest.py" "pypy perftest.py" -w 3 Benchmark #1: target/release/perftest Time (mean ± σ): 624.8 ms ± 9.8 ms [User: 623.0 ms, System: 0.8 ms] Range (min … max): 614.5 ms … 644.0 ms 10 runs Benchmark #2: php perftest.php Time (mean ± σ): 697.8 ms ± 18.3 ms [User: 696.7 ms, System: 1.1 ms] Range (min … max): 650.1 ms … 718.0 ms 10 runs Benchmark #3: python perftest.py Time (mean ± σ): 3.326 s ± 0.071 s [User: 3.313 s, System: 0.003 s] Range (min … max): 3.232 s … 3.419 s 10 runs Benchmark #4: pypy perftest.py Time (mean ± σ): 270.7 ms ± 5.7 ms [User: 257.5 ms, System: 13.0 ms] Range (min … max): 257.8 ms … 277.8 ms 11 runs Summary 'pypy perftest.py' ran 2.31 ± 0.06 times faster than 'target/release/perftest' 2.58 ± 0.09 times faster than 'php perftest.php' 12.29 ± 0.37 times faster than 'python perftest.py'
- andybak 6y ago[EDIT - posted in haste. I should RTFA] Ctrl+F javascript - nothing. At first glance this seems to be "dynamism==slow" but surely you need to explain why Python is slower than Javascript and for many years has resisted a lot of effort to match the performance of v8 and it's cousins?
- oneiftwo 6y agoI've always assumed array iteration is more expensive you don't know the size of the objects and/or they aren't contiguous.
- antb123 6y agohmm so complain and then say pypy is 7 times faster(and 4 times faster than nodejs)
- dguaraglia 6y agoI don't see a complaint, just an analysis of things that make code slow. Assuming that PyPy will just fix the problem is unrealistic, considering PyPy's limitations (doesn't support every architecture Python supports, lags behind a Python major minor version or two, etc.) Python is a great language and PyPy a great tool, but let's not become complacent or - worse - dismiss valid information just because we like them..
- FartyMcFarter 6y agoThis article seems to be using a very specific definition of interpreter, which is perhaps not what most people think of when they hear "interpreter" ? If I understand correctly, they call the module generating Python opcodes from Python code the "interpreter", and everything else is a "runtime". But Python opcodes are highly specific to CPython, and they are themselves interpreted, right? Calling the former "interpreter" and the latter something else seems like an artificial distinction. Not only is this definition of "interpreter" strange, but their definition of "runtime" also seems strange; in other languages, the runtime typically refers to code that assists in very specific operations (for example, garbage collection), not code that executes dynamically generated code.
- tom_mellior 6y ago> If I understand correctly, they call the module generating Python opcodes from Python code the "interpreter", No, you misunderstand. They explicitly define the interpreter as "ceval.c which evaluates and dispatches Python opcodes". Maybe "evaluate and dispatch" suggest something else to you, but ceval.c really is the code that iterates over a list of opcodes and executes the associated computations. This is absolutely 100% the part of Python that is the interpreter. The module that generates Python opcodes is the "compiler" (or "bytecode compiler"), and the article specifically points out that it's not included.
- gsnedders 6y agoBut at what point do function calls from ceval.c stop counting as part of the interpreter? Okay, calling a C function clearly at some point ceases being part of the interpreter, but is the entirety of an attribute lookup (i.e., executing the LOAD_ATTR inst) on an ordinary Python object part of the interpreter or the runtime? In plenty of VMs the object representation is an intrinsic part of the VM design, with the VM having deep knowledge of it.
- tom_mellior 6y agoThat's a very blurry distinction, and I'm not very interested in it. I was correcting the OP's first two paragraphs, I didn't take a stance on the third.
- eggsnbacon1 6y agofor reference, a pure java version through JMH that takes 0.38 seconds on my machine. This uses parallel stream so its multithreaded. Single threaded it takes 0.71 seconds. Removing the blackhole to allow dead code elimination takes single thread down to 0.41 seconds. This is close to PyPy, which I assume is dead code eliminating the string conversion as well. package org.example; import org.eclipse.collections.api.block.procedure.primitive.IntProcedure; import org.eclipse.collections.impl.list.Interval; import org.openjdk.jmh.annotations.Benchmark; import org.openjdk.jmh.infra.Blackhole; public class MyBenchmark { @Benchmark public void testMethod(final Blackhole blackhole) { Interval.oneTo(20) .parallelStream().forEach((IntProcedure) outer -> Interval.oneTo(1000000) .forEach( (IntProcedure) inner -> // prevent dead code elimination blackhole.consume(Integer.toString(inner)))); } }
- 6c696e7578 6y agoNot too surprising as dynamic languages take 4-20 times longer at numerical work, from rough experience. Java/c/c#/c++/rust etc are roughly at the same end of the spectrum (unless you're creating stupid numbers of objects). Perl/python/ruby, they're dynamic, so expect slower results. I like the threaded approach you're using.
- eggsnbacon1 6y ago> like the threaded approach you're using. thanks :) Eclipse Collections can also do batched loops which might speed this up. Telling Java how many threads you want will probably help as well Interestingly, I tried plain old for(i) loops and the result was exactly the same. At least for Eclipse Collections, the syntactic sugar for ranges and forEach is completely optimized away, apparently.
- palinkapika 6y agoYou might already know this, but there is one potential caveat with APIs like that when it comes to performance or at least measuring their performance. The hot loop (the code that actually iterates over your data points) does not live in your code base. But performance often depends on what the JIT compiler makes out of that loop. If the API is used at several locations in your program, the compiler might not be able to generate code that is optimal for your call-site and inputs. Instead, it will generate generalized code that works for all inputs but might be slower. However, when writing benchmarks, there is often no other code around to force the JIT compiler to generate such generalized code. The following code demonstrates this. If the parameter warmup is set, I invoke the forEach methods with different inputs first (they do the same but are different methods in the Java byte code). The purpose is to force the compiler to generate generalized code: @Param({"true", "false"}) public boolean warmup; @Setup public void setup() { if (!warmup) { Interval.oneTo(20).forEach( (IntProcedure) i -> Interval.oneTo(1_000_000).forEach( (IntProcedure) j -> Integer.toString(j))); Interval.oneTo(20).forEach( (IntProcedure) i -> Interval.oneTo(1_000_000).forEach( (IntProcedure) j -> Integer.toString(j))); Interval.oneTo(20).forEach( (IntProcedure) i -> Interval.oneTo(1_000_000).forEach( (IntProcedure) j -> Integer.toString(j))); } } @Benchmark public void coollectionsBlackhole(Blackhole blackhole) { Interval.oneTo(20).forEach( (IntProcedure) i -> Interval.oneTo(1_000_000).forEach( (IntProcedure)j -> blackhole.consume(Integer.toString(j)))); } @Benchmark public void collectionsDeadCode() { Interval.oneTo(20).forEach( (IntProcedure) i -> Interval.oneTo(1_000_000).forEach( (IntProcedure)j -> Integer.toString(j))); } And the for loop implementations for reference: @Benchmark public void loopBlackhole(Blackhole blackHole) { for (int j = 0; j < 20; j++) { for (int i = 0; i < 1_000_000; i++) { blackHole.consume(Integer.toString(i)); } } } @Benchmark public void loopDeadCode() { for (int j = 0; j < 20; j++) { for (int i = 0; i < 1_000_000; i++) { Integer.toString(i); } } } @Benchmark public void loopBlackholeOnly(Blackhole hole) { for (int j = 0; j < 20; j++) { for (int i = 0; i < 1_000_000; i++) { hole.consume(0xCAFEBABE); } } } On my desktop machine, this gives me the following results (Java 11/Hotspot/C2): Benchmark (warmup) Mode Cnt Score Error Units collectionsDeadCode true avgt 99 162.482 ± 1.615 ms/op collectionsDeadCode false avgt 99 218.816 ± 4.217 ms/op coollectionsBlackhole true avgt 99 235.122 ± 1.362 ms/op coollectionsBlackhole false avgt 99 270.192 ± 1.627 ms/op loopBlackhole true avgt 99 207.214 ± 1.162 ms/op loopBlackhole false avgt 99 206.711 ± 0.932 ms/op loopBlackholeOnly true avgt 99 74.774 ± 0.180 ms/op loopBlackholeOnly false avgt 99 74.359 ± 0.180 ms/op loopDeadCode true avgt 99 143.394 ± 0.900 ms/op loopDeadCode false avgt 99 142.654 ± 0.795 ms/op All results are for single threaded code. Now, the difference is not huge, but significant. Overall it looks like the invocation of toString(int) is not removed and accounts for most of the runtime. Just to be clear: I am not saying one should stay way from the stream APIs. As soon as the work per item is more than just a few arithmetic operations, there is a good chance the difference in runtime is negligible. But when doing numerical work (aggregations etc.), a simple loop might be the better option. Finally, these differences are of course compiler-dependent. For example, C2 might behave differently than Graal and who knows what the future brings.
- Mikhail_K 6y agoOn my somewhat old Linux machine his main() takes 5.22 seconds. Meanwhile, this Julia code @time map(string, 1:1000000); reports execution time 0.18 seconds. But that includes compilation time, if you use BenchmarkTools that runs the code repeatedly, I get 88.6 milliseconds
- eggsnbacon1 6y agothe outer loop runs 20 times as well
- umvi 6y agoI've found it often doesn't matter how fast or slow python is if the bottleneck is outside of python's control. For example, I wrote an lcov alternative in python called fastcov[0]. In a nutshell, it leverages gcov 9's ability to send json reports to stdout to generate a coverage report in parallel (utilizing all cores) Recently someone emailed me and advised that if I truly wanted speed, I needed to abandon python for a compiled language. I had to explain, however, that as far as I can tell, the current bottleneck isn't the python interpreter, but GCC's gcov. Python 3's JSON parser is fast enough that fastcov can parse and process a gcov JSON report before gcov can serialize the next JSON. So really, if I rewrote it in C++ using the most blisteringly fast JSON library I could find, it would just mean the program will spend more time blocking on gcov's output. In summary: profile your code to see where the bottlenecks are and then fix them. Python is "slow", yes, but often the bottlenecks are outside of python, so it doesn't matter anyway. [0] https://github.com/RPGillespie6/fastcov https://github.com/RPGillespie6/fastcov
- zeta0134 6y agoAlso of course, like any other language, Python is not immune to I/O wait. Recently I was using Python to do some low level audio processing (nothing fancy) and wondering why it was taking so long to process such a small file. Turns out the "wave" stream reader I was using isn't internally buffered, like most of the normal file readers are, so each single-byte read was making the round trip to disk. Simply reading the whole file in one go sped up the program's execution more than 10x. I like to think I've been doing this long enough to not make such silly mistakes, but when you're focused on a totally different area of the problem, things like this are still quite easy to miss.
- nh2 6y agoI'd argue that these things are comparatively easy to address, as one use of `strace` will immediately show bad IO patterns like that, and it works the same way across all programming languages.
- deleted 6y ago[deleted]
- stabbles 6y agoThis is a fun benchmark in C++, where you can see that GCC has a more restrictive small string optimization. On my desktop the main python example runs in 3.1s. Then this code void example() { for (int64_t i = 1; i <= 20; ++i) { for (int64_t j = 1; j <= 1'000'000; ++j) { std::to_string(j); } } } runs with GCC in 2.0s and with clang in 133ms, so 15x faster. I've also benchmarked it in Julia: function example() for j = 1:20, i = 1:1_000_000 string(i) end end which runs in 592ms. Julia has no small string optimization and does have proper unicode support by default. None of the compilers can see that the loop can be optimized out.
- eggsnbacon1 6y agoThe Clang result is suspiciously fast. Way faster than GCC and my Java attempt. Maybe its dead code pruning the string conversion? Java was about the same as Julia for this, on my machine
- runevault 6y agoIf it's using string optimizations due to short strings, you're cutting allocations in almost half I believe, as well as putting said strings on the stack instead of the heap which will give... significant performance gains. Short strings in the C++ standard library can be crazy fast.
- plorkyeran 6y agoIt doesn't merely cut allocations in half; it cuts them to zero.
- runevault 6y agoDo you mean allocations as in heap? If so yeah, I guess I should have stated it better as it still needs to put space on the stack for the short string.
- 6y ago
- apalmer 6y agoI think he is making the distinction between 3 different categories that readers are in general lumping into the 'interpreter': 1) Python is executed by an interpreter (necessary overhead) 2) Python as a language is so dynamic/flexible/erfonomic that it has to do things that have overhead (necessary complexity unless you change the language) 3) the specific implementation of the interpreter achieves 1 and 2 in ways that can be significantly slower than necessary Seems he is pointing out that a lot of performance issues that are generally thought to be due to 1 and 2 are really 3
- gsnedders 6y ago> 2) Python as a language is so dynamic/flexible/erfonomic that it has to do things that have overhead (necessary complexity unless you change the language) As PyPy demonstrates, much of this doesn't need to have anywhere near the overhead that it does in CPython. You can absolutely do better than CPython without changing the language, and you can do better than CPython without a JIT if you start specialising code.
- varelaz 6y agoI really sick of such kind of benchmarks. I never have seen a real world python program that doesn't depend on any IO or doesn't have some C code behind wrapped calls. If your code is heavy with computations there are numpy/scipy libs that are very good at this. These optimizations bring < 10% of speed to real project/programm, but will require a lot of developers time to support it. If performance is the key feature and very critical, then likely python is not the right choice, because python is more about flexibility, ability to maintain and write solid, easy to read code.
- VHRanger 6y agoHard disagree. Learning the tool you're working with means you know patterns to write generally more efficient code. Even if you're going to use numpy/cython/cffi for faster submodules, writing faster code in general is a good thing.
- varelaz 6y agoI don't mind about knowing limitations. I'm saying that these optimizations usually are very hard tradeoffs and opposite side of it – code readability, speed of developer work, ability to maintain it working. I tried Cython and PyPy. Both are really good if your project is started with them, but if you decide to migrate to them in order to increase performance, it's like rewrite project to another language. Also both have a lot of limitations and cpython gives you still a lot more flexibility in decision making (choose frameworks, libraries and approaches how to solve certain problem)
- optimuspaul 6y agoI fail to see how you actually disagree.
- vmchale 6y agoPeople write web services in Python.
- sandGorgon 6y ago>impressively, PyPy [PyPy3 v7.3.1] only takes 0.31s to run the original version of the benchmark. So not only do they get rid of most of the overheads, but they are significantly faster at unicode conversion as well. It is super unfortunate that Pypy struggles for funding every quarter. Funding a mil for Pypy by the top 3 python shops (Google, Dropbox, Instagram?) should be rounding error for these guys...and has the potential to pay off in hundreds of millions atleast (given the overall infrastructure spend).
- staticassertion 6y agohttps://blog.pyston.org/ https://blog.pyston.org/ You can look at the last release of Pyston for why companies are not funding this work. They're just moving to other languages. Consider that Dropbox, Google, and Instragram likely spend much more than 1M on optimizing/ improving Python, just to have it still be a relatively slow language. At some point it becomes way cheaper to just move to other languages that don't require that constant level of effort. Think about how much time and money is spent on things like: * Adding types to Python * Improving performance when you could just move performance critical code to Go/ Rust as Dropbox has done, for a fraction of the cost, with far less maintenance burden. "Make Python faster" is just a losing game imo. It is fundamentally never going to be as fast as other languages, it's far too dynamic (and that's a huge appeal of the language). Just look at the optimizations done here - moving `str` to local scope to avoid a global lookup? And can you even avoid that? Not without changing semantics - what if I want to change `str` globally? Still, I was surprised by some of the sorta nutty wins that were achieved here. There's clearly some perf left on the table, and I'm not an expert on interpreters, it just seems really hard to build a semantically equivalent Python ("I can change global functions from a background thread") that can automatically optimize these things.
- disgruntledphd2 6y agoRewrites are the devil though, at least that's what Joel Spolsky taught me. Additionally, spending time at a really large tech company taught me that rewrites apparently never work (or at least it's better to build compilers for your shitty slow code rather than rewrite it in a better language). The linked post concerns Dropbox, who to be fair are an infrastructure company with low margins selling a commodity product. It doesn't surprise me that they would find such a move more important than say, Instagram who have a very large margin selling a differentiated product. I do really appreciate python types (and things like monkeypatch from IG for adding them) as they are an incremental improvement to a widely used language.
- earthboundkid 6y agoI would be interested in seeing the performance difference for f"{i}". My intuition is that it would be faster.
- hpcjoe 6y agoI know this is supposed to be about python optimization. However, the post switches over to C at nearly the beginning of the process. Hence it really is about how to optimize python applications, by rewriting them in C (or other fast languages). Which IMO, isn't about optimizing Python, apart from tangential API/library usage. I've been under the impression that I'll get the best performance out of a language when I write code that leverages the best idiomatic features, and natural aspects of the language. If I have to resort to another language along the way to get needed performance, I guess the question is, isn't this a strong signal that I should be looking at different languages ... specifically fast ones? Most of the fast elements of python are interfaces to C/Fortran code (numpy, pandas, ...). What is the rationale for using a slow language as glue versus using a faster language for processing?
- shepardrtc 6y ago> What is the rationale for using a slow language as glue versus using a faster language for processing? It's much quicker and easier to put together a Python program than a C/C++ program. The number of libraries out there for Python is incredible.
- hpcjoe 6y agoHonestly, this is subjective. I generally agree that some languages are well designed for fast prototyping, but not really great for computationally intensive production. I see Python in this role. Note that when the original author wished to move this to a faster execution capability, they had to change languages. So they get the disadvantage of having to do the port anyway. How does this save time?
- jokoon 6y agoI'll never understand why there are so many fast, alternative python interpreters. Is the language correctness of the official interpreter causing lower performance? Does it prevent using some modules? What's at stake here? I'm planning to use python as a game scripting language but I hear so much about performance issues that it scares me to try learning how to use it in a project. I love python though.
- entha_saava 6y agoThey are mostly JITs or have specific purposes (like stackless python). JITs like pypy have a raw performance advantage. But implementation is more complex, and there is startup time overhead. Also I have heard CPython reference implementation prefers clarity of code over optimizations. As for game scripting, Lua is most used one.
- inglor 6y agoThe article is quite good but the Node.js example is wrong. OP is measuring dead code elimination and general overhead time. No strings get harmed in the process, run Node.js with --trace-opt to see what's happening.
- kmod 6y agoI don't know how to interpret the results of --trace-opt, but when I increase the iteration bounds 10x the running time increases 10x, so I don't think the code is being eliminated. This is with node v12.13.0
- erdewit 6y agoReplacing str(i) with the f-string f'{i}' lets it run about 2x faster.
- dzonga 6y agothis is an area, nim could've come and improved on. i.e if nim had an ability to port your python code and have it running 98% on nim. most python users would have been there already.
- deleted 6y ago[deleted]
- rurban 6y agoI did similar studies about a decode ago for perl with similar results. But what he's missing are two much more important thing. 1. smaller datastructures. They are way overblown, both the ops and the data. Compress, trade for smaller data and more simplier ops. In my latest VM I use 32 bit words for each and for each data. 2. Inlining. A tellsign is when the calling convention (arg copying) is your biggest profiler contributor. Python's bytecode and optimizer is now much better than perl's, but it's still 2x slower than perl. Python has by far the slowest VM. All the object and method hooks are insane. Still no unboxing or optimize refcounting away, which brought php ahead of the game.
- tuananh 6y agoCan anyone explain to me why this is a lot faster than j.toString() or String(j) for (let i = 0; i < 20; i++) { for (let j = 0; j < 1000000; j++) { `${j}` } } I got Executed in 177.12 millis fish external usr time 159.12 millis 91.00 micros 159.03 millis sys time 17.96 millis 443.00 micros 17.52 millis
- aldanor 6y agoSince people are talking about speed and JIT here, it's worth mentioning Numba (http://numba.pydata.org http://numba.pydata.org). Being in the quant field myself, it's often been a lifesaver - you can implement a c-like algorithm in a notebook in a matter of seconds, parallelise it if needed and get full numpy api for free. Often times if you're doing something very specific, you can beat pandas/numpy versions of the same thing by an order of magnitude.
- adsharma 6y agoStatic subset of <language> While many here correctly observe that the "too much dynamism" can become a performance bottleneck, one has to analyze the common subset of python that most people use and see how much dynamism is intrinsic to those use cases. Other languages like JS have tried a static subset (Static typescript is a good example), that can be compiled to other languages - usually C. Python has had RPython, but no one uses it outside of the compiler community. The argument here is that python doesn't have to be one language. It could be 2-3 with similar syntax catering to different use cases and having different linters. * A highly dynamic language that caters to "do this task in 10 mins of coding" use case. This could be used by data scientists and other data exploration use cases. * A static subset where performance is at a premium. Typically compiled down to another language. Strict typing is necessary. Performance sensitive and a large code base that lives for many years. * Some combination of the two (say for a template engine type use case). A problem with the static use case is that the typing system in python is incomplete. It doesn't have pattern matching and other tools needed to support algebraic data types. Newer languages such as swift, rust and kotlin are more competitive in this space.
- deleted 6y ago[deleted]