4 ms·
I favorited this post to read the discussion later, but looking into that "plan", it's mostly a bunch of nothing. It's very hard to estimate how much speedup a
by latenightcoding 6y ago
I favorited this post to read the discussion later, but looking into that "plan", it's mostly a bunch of nothing.
It's very hard to estimate how much speedup a JIT will get you on a dynamic language like python and x5 speedup seems unrealistic.
There are other lower hanging fruits, like optimizing core data structures (e.g: the implementation of python dicts) .
- cuchoi 6y agoWhat makes you think that the implementation of dicts is a low hanging fruit? Only asking because there has been a lot of work to make them faster: https://www.youtube.com/watch?v=npw4s1QTmPg https://www.youtube.com/watch?v=npw4s1QTmPg
- jammycrisp 6y agoAgreed on optimizing core objects. I recently wrote a C base class (https://jcristharif.com/quickle/#structs-and-enums https://jcristharif.com/quickle/#structs-and-enums) for defining dataclass-like-types that's noticeably faster (~5-10x) to init/copy/serialize/compare than other options (dataclasses, pydantic, namedtuples...). For some applications I write this has a non-negligible performance impact, without requiring deep interpreter changes. Using the base class is nice - my application objects are still defined in normal python code, but all the heavy lifting is done in the c-extension. However, this speedup comes at the cost of being less dynamic. I'm not sure how much more optimized core python objects could be without sacrificing some of the dynamism some programs rely on. Python dicts are already pretty optimized as is.
- byko3y 6y ago>I recently wrote a C base class (https://jcristharif.com/quickle/#structs-and-enums https://jcristharif.com/quickle/#structs-and-enums) for defining dataclass-like-types that's noticeably faster (~5-10x) to init/copy/serialize/compare than other options (dataclasses, pydantic, namedtuples...) YouTube also encountered the same problem. Their solution sounds kinda like "never use pickle, because it's slow. Use custom serialization".
- metalliqaz 6y agoI'm not saying you're wrong, but last I checked (admittedly, years ago) there was at least 1 Python JIT with a lot of development and user hours behind it, with good results. Seems like that experience should yield some decent estimates on speedup.
- d0mine 6y ago> x5 speedup seems unrealistic. I see x4 speedup https://speed.pypy.org https://speed.pypy.org
- chubot 6y agoYeah the initial page looked very thin, but this linked one has a tiny bit more info: https://github.com/markshannon/faster-cpython/blob/master/tiers.md https://github.com/markshannon/faster-cpython/blob/master/ti... Still I'm a bit skeptical ... his reasons for why others failed and he will succeed is not that convincing.
- formerly_proven 6y agofwiw CPython is a really naive bytecode interpreter. Way more naive and basic than people seem to generally assume.
- FridgeSeal 6y agoIt also does (to my knowledge) no real optimisation on user code. I've also read that a lot of pythons construction and language semantics make it difficult to implement more performance optimisations.
- coldtea 6y ago>It's very hard to estimate how much speedup a JIT will get you on a dynamic language like python and x5 speedup seems unrealistic. You can have a 2x or 10x for many use cases speedup without a JIT, as PHP7 proved. You just need to start with a slow, not very optimized, implementation, which CPython pretty much is. As for 5x, Javascript has had much more than speed bump than that with its JITs (compared to the interpreted Javascript pre-JSCore, Tracemonkey and V8 circa 1997-2005) and it's just as dynamic as Python... >There are other lower hanging fruits, like optimizing core data structures (e.g: the implementation of python dicts). Funny that you should mention it, because the author of the proposal has already done significant work (available since Python 3.3. or so) optimizing the dicts...
- imtringued 6y agoFor all I know, Python is extremely slow. If you think a 5x improvement is not possible then you may be underestimating how slow Python is.
- viraptor 6y agoOr aware how dynamic it is. The ability to run 'locals()' for example makes a lot of possible improvements quite tricky / more heavy than necessary.
- kzrdude 6y agopypy has been working on it, for 10 years though, and their average is 4x on their benchmarks. Sure, I think this means they do 10x or 20x performance on certain tasks, and that's impressive. But still, it's not easy to do 5x across the board.
- evilotto 6y agoIs there any comprehensive overview of why python is slow in the first place? There seems to be opportunities for small optimizations, but is that the only place where the slowness is? A few years ago I tried to find out what investigations had been done on the speed of bytecode generation and there was basically nothing. No one could even suggest how long bytecode generation took, only that "it's really slow". Isn't the common wisdom to measure, then optimize, not the other way around?
- mjburgess 6y agoIt has nothing to do with it being interpreted (this is a common misidentified issue with dynamic langs). The fundamental issue is that python is a pointer machine: everything requires a dynamic lookup in memory. Eg., x = [1, 2, 3] len(x) Here `x` is an actual string in memory which is a key in a locals() dictionary which holds values. (cf. with C where it is just a memory address). Likewise the list is a list of pointers (not a sequential array). And its heterogenous, ie., the contents can be of any type. Likewise `len` is a string into a dictionary of functions which has to be looked up. etc. The whole thing is many levels of indirection. Applying an operation to a value (eg., even x + y) requires jumping around the memory of the machine many times. This is necessary, in general, to deliver on the dynamic lang. features python provides. Julia solves some of these issues by using static type information to ditch this dynamic behaviour. My suspicion is that python can follow a similar path (eg., above, x should be compiled to a static homogenous array of ints).
- jashmatthews 6y agoLocal variables in CPython are stored in the stack frame and accessed via LOAD_FAST and STORE_FAST without doing anything with x as a string or a hash table: http://stupidpythonideas.blogspot.com/2015/12/how-lookup-works.html http://stupidpythonideas.blogspot.com/2015/12/how-lookup-wor...
- mjburgess 6y agoMy claim there was only that it was reified in memory, not that in the case of loading x, you had to use it. Though it is worth noting that you dont. In general the point stands, the reason for slowness is indirection & reification. (Not sure why i'm downvoted).
- byko3y 6y agoFrankly speaking, I had this impression too. The reason why JVM and JS uses own bytecode for implementing JIT-compilation is because main means of optimization are inlining and vector scalarization (convert complex object to simple ones). E.g. you have a container that hold integer and does nothing else, thus when JIT-compiling you can skip the whole boxing/unboxing and manipulate a single integer value stored in register. However, you cannot do that in a simple ways if your integer container is a C extension black box — you cannot "inline" the container's functions and reduce its store-loads. This is why PyPy reimplements standard library in RPython — so you can JIT-optimize it. But it feels like Mark Shannon knows nothing about these efforts — which is kinda strange considering his position of core CPython developer. >There are other lower hanging fruits, like optimizing core data structures (e.g: the implementation of python dicts) Unfortunately, you cannot easily implement efficient data containers without rewriting existent python code. The latter one relies heavily on dictionary-based access to pretty much everything, and you cannot easily convert "string hash" access into "record offset" access, because you cannot know a priori what object has what structure and converting hash into offset is basically the same dictionary lookup. For example: a = A() a.field = varname + 1 What can you optimize here? What "varname" is? What A's structure is? Is "A" a class or a function? Not only you are unable tell the semantic of the code just by looking at the code — you can't even tell the semantic after you've examined the "A" and "varname" on some previous iteration, because somebody might've declared/modified those on outer scope or directly modified "A" or "varname". Last year in my spare time I've been working on an unpublished library for python multitasking with shared memory structures (probably will make some blog post in few weeks and link it here), and I also encountered the problem of inherently inefficient implementation of python basic types. However, I'm yet to find the solution without breaking compatibility with existing code. For example, if you look at ctypes, they have some very efficient containers, but using them in a regular python code is a pain, and the c-python interface eats most performance benefits of efficient containers. So what's really needed for optimization of python is some kind of python subset, like RPython but probably more human-friendly, so efficient containers can really become efficient while automatic optimizer can select or create automatically those efficient containers. Just like V8 JS engine does, which stores objects in records with static structure. It happens to works in JS for most cases. Countrary, in Python it does not work for most cases, that's why we have so much struggle optimizing the Python.