5 ms·
I have to disagree that those are "interpreter" overheads, since the later versions of the benchmark do not use the Python interpreter at all yet still suffer f
by kmod 6y ago
I have to disagree that those are "interpreter" overheads, since the later versions of the benchmark do not use the Python interpreter at all yet still suffer from these overheads. Maybe disagreeing on the wording is pedantic though, since I think the real discussion is what can be done about it.
We definitely agree directionally: you can create specific implementations of operations that are fast for their inputs. But look at how specific we have to get here. It's not just on the type of the callable: we have to know the exact callable we are calling (the str type). And its behavior is heavily dependent on its argument type. So we need a code path that is specifically for calling the str() function on ints. I would argue that this is prohibitively specialized for an interpreter, but one can imagine a JIT that is able to produce this kind of specialization, and that's exactly what this blog post is trying to motivate.
- acqq 6y ago> So we need a code path that is specifically for calling the str() function on ints. I would argue that this is prohibitively specialized for an interpreter, I would argue it isn't. It's actually less about the ints but about the tuple allocation: "Now that we've removed enough code we can finally make another big optimization: no longer allocating the args tuple This cuts runtime down to 1.15s." (down from 1.40s) It seems to me that having special case for one argument, avoiding the tuple allocation in each call is nothing "prohibitively expensive" and would benefit all the functions being called with one argument, and there are enough of such. Regarding ints vs something else -- somewhere in the code there is anyway different code that does str of int vs str of something else. It's just about accessing that code for that type, it doesn't have to cost having today's speculative execution in the CPU's. The people who produce compiled code know about that and can use it.
- kmod 6y agoThere's a minor but important point: the amount of work str() has to do depends on its input type. In particular, if arg.__str__() returns a string or a subtype of string or not a string, str() will call different tp_init functions. So we can't get rid of the args tuple until after the optimization of knowing that arg.__str__() returns an string (which is why I did them in that order). So I guess the condition is a little bit weaker than str(int) but it's more than str() of one arg.
- burfog 6y agoI think you've proven that the language is a lost cause. It's never going to perform OK. With all the breakage of Python 3, it's a shame that the causes of slowness were not eliminated. The language could have been changed to allow full ahead-of-time compilation, without all the crazy object look-up. Performance might have been like that of Go.
- khc 6y agothen it will become perl6 which no one uses
- lizmat 6y agoPerl 6 is mostly not used because it has been renamed to Raku (https://raku.org https://raku.org using the #rakulang tag on social media). Raku is very much being used. If you want to stay up-to-date with Raku developments, you can check out the Rakudo Weekly News https://rakudoweekly.blog https://rakudoweekly.blog
- acqq 6y agoI'm still very sure that the Python's source code can be refactored to avoid the allocation completely in the case of a zero or one parameter for all the functions that now always have an expectation of the tuple. But I also understand that you can't achieve that in your case without changing the existing code base.
- whizzter 6y agoIt's all tied together because there is an shared value/calling model that permeates both runtime and interpreter and imo with CPython much of it becomes academic because unless people are willing to break source compatibility with native extensions(Esp with Python2 vs Python3 still lingering even if most have gone to P3) it's too much of a historic mess to clean up (Although it's a good sign that the official Python docs points to CFFI and ctypes and discourages new development with the C api). After doing my thesis work(JS-subset -> C compiler) i was a bit disappointed in the results on weaker processors and it led me to experiment with a series of toy compilers. Those experiments really showed how instrinsic a good value model is to determining your performance _ceiling_. From my memory you had a factor 3-5 of improving or eliminating each of these: - bytecode dispatch (vs native code) - naive dynamic type tests(vs optimized paths) - obj/refcount management (vs just using values directly with a GC backed model with cheap stack-root-registrations) - naive property lookups(vs optimized variants). If you started with a naive dynamic js/python-like language ref counted impl you were within a magnitude of the CPython level of perf, as you check off these checkboxes you get closed to V8 (I used C compilers as backends for code generation so i had a slight edge on the lowlevel code generation part). The CPython model is great for integrating with hand written C/C++ code but the price you pay is a low ceiling since it requires you to manage ref counts (at the very least at function call boundaries even if you optimize them with deferrals) and as you mention in the article pass values via objects.
- barrkel 6y agoSpecial str for ints is not unusual, I'll specifically disagree with you there. Formatting and string output are incredibly common, especially in scripting languages, and repay specialized optimization. IIRC, the Delphi compiler converts WriteLn('foo: ', 42) into WriteStr, WriteInt and WriteLn calls, for example. StringBuilder classes in Java and .NET have overloads to specialize appending non-boxed int. Etc.