9 ms·
Python 3.15’s interpreter for Windows x86-64 should hopefully be 15% faster
- machinationu 10mo agoThe Python interpreter core loop sounds like the perfect problem for AlphaEvolve. Or it's open source equivalent OpenEvolve if DeepMind doesn't want to speed up Python for the competition.
- g947o 10mo ago> This has caused many issues for compilers in the past, too many to list in fact. I have a EuroPython 2025 talk about this. Looks like it refers to this: https://youtu.be/pUj32SF94Zw https://youtu.be/pUj32SF94Zw (wish it's a link in the article)
- eru 10mo ago> (wish it's a link in the article) I've asked Ken. He said he'll update the article.
- Hendrikto 10mo agoTLDR: The tail-calling interpreter is slightly faster than computed goto. > I used to believe the the tailcalling interpreters get their speedup from better register use. While I still believe that now, I suspect that is not the main reason for speedups in CPython. > My main guess now is that tail calling resets compiler heuristics to sane levels, so that compilers can do their jobs. > Let me show an example, at the time of writing, CPython 3.15’s interpreter loop is around 12k lines of C code. That’s 12k lines in a single function for the switch-case and computed goto interpreter. > […] In short, this overly large function breaks a lot of compiler heuristics. > One of the most beneficial optimisations is inlining. In the past, we’ve found that compilers sometimes straight up refuse to inline even the simplest of functions in that 12k loc eval loop.
- kccqzy 10mo agoI think in the protobuf example the musttail did in fact benefit from better register use. All the functions are called with the same arguments, so there is no need to shuffle the registers. The same six register-passed arguments are reused from one function to the next.
- cma 10mo agoDoes MSVC support computed goto?
- mishrapravin441 10mo agoReally nice results on MSVC. The idea that tail calls effectively reset compiler heuristics and unblock inlining is pretty convincing. One thing that worries me though is the reliance on undocumented MSVC behavior — if this becomes widely shipped, CPython could end up depending on optimizer guarantees that aren’t actually stable. Curious how you’re thinking about long-term maintainability and the impact on debugging/profiling.
- kenjin4096 10mo agoThanks for reading! For now, we maintain all 3 of the interpreters in CPython. We don't plan to remove the other interpreters anytime soon, probably never. If MSVC breaks the tail calling interpreter, we'll just go back to building and distributing the switch-case interpreter. Windows binaries will be slower again, but such is life :(. Also the interpreter loop's dispatch is autogenerated and can be selected via configure flags. So there's almost no additional maintenance overhead. The main burden is the MSVC-specific changes we needed to get this working (amounting to a few hundred lines of code). > Impact on debugging/profiling I don't think there should be any, at least for Windows. Though I can't say for certain.
- mishrapravin441 10mo agoThat makes sense, thanks for the detailed clarification. Having the switch-case interpreter as a fallback and keeping the dispatch autogenerated definitely reduces the long-term risk.
- pxeger1 10mo agoProfile of llm generated comments
- mishrapravin441 10mo agoust to clarify, I’m writing these comments myself. I use grammar llm plugin though to clean up phrasing, but the substance is mine.
- develatio 10mo agoif the author of this blog reads this: can we can an RSS, please?
- kenjin4096 10mo agoGot it. I'll try to set one up this weekend.
- develatio 10mo agoThank you so much!!
- redox99 10mo agoThis seems like very low hanging fruit. How is the core loop not already hyper optimized? I'd have expected it to be hand rolled assembly for the major ISAs, with a C backup for less common ones. How much energy has been wasted worldwide because of a relatively unoptimized interpreter?
- deleted 10mo ago[deleted]
- kccqzy 10mo agoPython’s goal is never really to be fast. If that were its goal, it would’ve had a JIT long ago instead of toying with optimizing the interpreter. Guido prioritized code simplicity over speed. A lot of speed improvements including the JIT (PEP 744 – JIT Compilation) came about after he stepped down.
- davidkhess 10mo agoShould probably mention that Guido ended up on the team working on a pretty credible JIT effort. Though Microsoft subsequently threw a wrench in it with layoffs. Not sure the status now.
- IshKebab 10mo agoIf performance was a goal... hell if it was even a consideration then the language would be very different.
- eru 10mo agoYour are mixing up eras. For comparison: when Javascript was first designed, performance wasn't a goal. Later on, people who had performance as a goal worked on Javascript implementations. Thanks to heroic efforts, nowadays Javascript is one of the language with decently fast implementation around. The base design of the language hasn't changed much (though how people use it might have changed a bit). Python could do something similar.
- mhh__ 10mo ago
- mananaysiempre 10mo agoThe money shot (wish this were included in the blog post): # if defined(_MSC_VER) && !defined(__clang__) # define Py_MUSTTAIL [[msvc::musttail]] # define Py_PRESERVE_NONE_CC __preserve_none # else # define Py_MUSTTAIL __attribute__((musttail)) # define Py_PRESERVE_NONE_CC __attribute__((preserve_none)) # endif https://github.com/python/cpython/pull/143068/files#diff-45baf725df91ed7826458cda8a17c2b4a2e5296504de1ed6a1c5a9ebe6390a47 https://github.com/python/cpython/pull/143068/files#diff-45b... Apparently(?) this also needs to be attached to the function declarator and does not work as a function specifier: `static void *__preserve_none slowpath();` and not `__preserve_none static void *slowpath();` (unlike GCC attribute syntax, which tends to be fairly gung-ho about this sort of thing, sometimes with confusing results). Yay to getting undocumented MSVC features disclosed if Microsoft thinks you’re important enough :/
- deleted 10mo ago[deleted]
- publicdebates 10mo agoImportant enough, or benefits them directly? I have no good guesses how improving Python's performance would benefit them, but I would guess that's the real reason.
- HPsquared 10mo agoI wonder if this is related to Python in Excel. You'll have lots of people running numerical stuff written in Python, running on Microsoft servers.
- mkoubaa 10mo agoA lot of commercial engineering and scientific software runs on windows.
- andix 10mo agoI guess there are some Python workloads on Azure, Microsoft provides a lot of data analysis and LLM tools as a service (not paid by CPU minutes). Saving CPU cycles there directly translates to financial savings.
- Rakshath_1 10mo ago[dead]
- bgwalter 10mo agoMSVC mostly generates slower code than gcc/clang, so maybe this trick reduces the gap.
- metaltyphoon 10mo agoIs this backed by real evidence?
- bluecalm 10mo agoMy experience is 10%-15% slower than GCC. That was 10 years ago though.
- pjmlp 10mo agoImagine how much fast those Windows and XBox games would be if they used gcc/clang. /s
- jtrn 10mo agoIm a bit out of the loop with this, but hope its not like that time with python 3.14, when it was claimed a geometric mean speedup of about 9-15% over the standard interpreter when built with Clang 19. It turned out the results were inflated due to a bug in LLVM 19 that prevented proper "tail duplication" optimization in the baseline interpreter's dispatch loop. Actual gains was aprox 4%. Edit: Read through it and have come to the conclusion that the post is 100% OK and properly framed: He explicitly says his approach is to "sharing early and making a fool of myself," prioritizing transparency and rapid iteration over ironclad verification upfront. One could make an argument that he should have cross-compiler checks, independent audits, or delayed announcements until results are bulletproof across all platforms. But given that he is 100% transparent with his thinking and how he works, it's all good in the hood.
- kenjin4096 10mo agoThanks :), that was indeed my intention. I think the previous 3.14 mistake was actually a good one on hindsight, because if I didn't publicize our work early, I wouldn't have caught the attention of Nelson. Nelson also probably wouldn't have spent one month digging into the Clang 19 bug. This also meant the bug wouldn't have been caught in the betas, and might've been out with the actual release, which would have been way worse. So this was all a happy accident on hindsight that I'm grateful for as it means overall CPython still benefited! Also this time, I'm pretty confident because there are two perf improvements here: the dispatch logic, and the inlining. MSVC can actually convert switch-case interpreters to threaded code automatically if some conditions are met [1]. However, it does not seem to do that for the current CPython interpreter. In this case, I suspect the CPython interpreter loop is just too complicated to meet those conditions. The key point also that we would be relying on MSVC again to do its magic, but this tail calling approach gives more control to the writers of the C code. The inlining is pretty much impossible to convince MSVC to do except with `__forceinline` or changing things to use macros [2]. However, we don't just mark every function as forceinline in CPython as it might negatively affect other compilers. [1]: https://github.com/faster-cpython/ideas/issues/183 https://github.com/faster-cpython/ideas/issues/183 [2]: https://github.com/python/cpython/issues/121263 https://github.com/python/cpython/issues/121263
- acemarke 10mo agoI've never seen this kind of benchmark graph before, and it looks really neat! How was this generated? What tool was used for the benchmarks? (I actually spent most of Sep/Oct working on optimizing the Immer JS immutable update library, and used a benchmarking tool called `mitata`, so I was doing a lot of this same kind of work: https://github.com/immerjs/immer/pull/1183 https://github.com/immerjs/immer/pull/1183 . Would love to add some new tools to my repertoire here!)
- eesmith 10mo agoAre you referring to the violin plot? https://en.wikipedia.org/wiki/Violin_plot https://en.wikipedia.org/wiki/Violin_plot and in Matplotlib as https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.violinplot.html https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.... It's in essence a histogram for the distribution, with smoothing, and mirrored on each side. It looks nice, but is not without well-deserved opposition because 1) the use of smoothing can hide the actual distribution, 2) mirroring contains no extra information, while taking up space, and implying the extra space contains information, and 3) when shown vertically, too often causes people to exclaim it looks like a vulva. In an HN discussion on the topic, medstrom at https://news.ycombinator.com/item?id=40766519 https://news.ycombinator.com/item?id=40766519 points to a half-violin plot at https://miro.medium.com/v2/1*J3Q4JKXa9WwJHtNaXRu-kQ.jpeg https://miro.medium.com/v2/1*J3Q4JKXa9WwJHtNaXRu-kQ.jpeg with the histogram on the left, and the half-violin on the right, which gives you a chance to see side-by-side presentation of the same data.
- Tarq0n 10mo agoHistograms aren't necessarily a true depiction of the distribution. Bin count or width has a large impact on what details get shown.
- eesmith 10mo agoSure. Very few distributions have lovely square edges, which otherwise indicate some very high frequencies in the distribution, or quantized values. But that also means we are used to seeing histograms and their bin count and widths in order to estimate possible variances from the true distribution;. While it's much harder to do the same with violin plots.
- deleted 10mo ago[deleted]
- forrestthewoods 10mo agoIs there a Clang based build for Windows? I’ve been slowly moving my Windows builds from MSVC to Clang. Which still uses the Microsoft STL implementation. So far I think using clang instead of MSVC compiler is a strict win? Not a huge difference mind you. But a win nonetheless.
- gozzoo 10mo agoI have quetion - slightly off topic, but related. I was wandering why is pyhton interpreter so much slower than V8 javascript interpreter when both javascript and python are dynamic interpreted languages.
- bheadmaster 10mo agoI can think of two possible reasons: First is the Google's manpower. Google somehow succeeds in writing fast software. Most Google products I use are fast in contrast to the rest of the ecosystem. It's possible that Google simply did a better job. The second is CPython legacy. There are faster implementations of Python that completely implement the API (PyPy comes to mind), but there's a huge ecosystem of C extensions written with CPython bindings, which make it virtually impossible to break compatibility. It is possible that this legacy prevents many possible optimizations. On the other hand, V8 only needs to keep compatibility on code-level, which allows them to practically switch out the whole inside in incremental search for a faster version. I might be wrong, so take what I said with a grain of salt.
- canucker2016 10mo agoDon't forget that there was a Google attempt at making a faster Python - Unladen Swallow. It got lots of PR but never merged with mainline CPython (wikipedia says a dev branch was released). see https://en.wikipedia.org/wiki/Unladen_Swallow https://en.wikipedia.org/wiki/Unladen_Swallow
- pansa2 10mo agoUnladen Swallow got a lot of hype but was only a very small project. IIRC the only people working on it were two interns. V8 was a much higher priority - Google hired many of the world’s best VM engineers to develop it.
- pjmlp 10mo agoSome of them like Lars Bak, have background up to Self VM, which is a language much more dynamic than Python. Anything goes regarding changing object shapes, it is one step further than Smalltalk in language plasticity.
- Quitschquat 10mo agoTbh, 15% faster than slow AF is still slow AF
- dingdingdang 10mo agoYup, but 5 to 15% faster year on year is real progress and that's ultimately what the big user base of Python are counting on at this point.. and they seem to be getting it! Full disclaimer: I'm not a heavy Python user exactly due to the performance and build/distribution situation - it's just sad from a user-end perspective (I'm not addressing centralised web deployment here but rather decentralised distribution which I ultimately find more "real" and rewarding).
- horizion2025 10mo agoI don't understand this focus on micro performance details... considering that all of this is about an interpretation approach which is always going to be slow relatively speaking. The big speed up would be to JIT it all, then you dont need to care about structuring of switch loops etc
- int_19h 10mo agoYou'd be surprised at how little speedup you get from simply JIT-compiling the Python bytecode. It's so high-level that most interesting stuff happens in the layers below anyway.
- horizion2025 10mo agoBut if that is so why this focus on the few clock cycles of dispatch?
- int_19h 10mo agoBecause it is a fairly easy thing - it's a code transform that's mostly mechanical. And it also improves code quality, unusual for an optimization. So if that nets you those extra few percent, why not?
- eab- 10mo agoMy understanding is that also this tail call based interpretation is also kinder to the branch predictor. I wonder if this explains some of the slow downs - they trigger specific cases that cause lots of branch mispredictions.
- DrewADesign 10mo agoAfter years of admonition discouraging me, I’m using Python for a Windows GUI app over my usual C#/MAUI. I’m much more familiar with Python and the whole VS ecosystem is just so heavy for lightweight tasks. I started with tkinter but found it super clunky for interactions I needed heavily, like on field change, but learning QT seemed like more of a lift than I was interested in. (Maybe a skill issue on both fronts?) Grabbed wxglade and drag-and-dropped an interface with wxpython that only has one external dependency installable with pip, is way more convenient than writing xaml by hand, and ergonomically feels pretty pythonic compared to QT. Glad to see more work going into the windows runtime because I’ll probably be leaning on it more.
- halfcat 10mo agoWait until you see ImGui bindings for Python [1]. It’s immediate mode instead of retained mode like Tkinter/Qt/Wx. It might not be what you’d want if you’re shipping a thick client to customers, but for internal tooling it’s awesome. imgui.text(f"Counter = {counter}") if imgui.button("increment counter"): counter += 1 _, name = imgui.input_text("Your name?", name) imgui.text(f"Hello {name}!") [1] https://github.com/pthom/imgui_bundle https://github.com/pthom/imgui_bundle
- DrewADesign 10mo agoThis looks like it would be perfect for the internal user that really just needs to run a shell script with options who’s in the “technical enough to follow instructions faithfully, not technical enough to comfortably/reliably use the command line” demographic.
- stinos 10mo agoImGui has been on my watchlist for years and recently I finally had an application which seemed I could put it to use. It essentially delivered on all points I hoped it would. After decades in software, it doesn't happen often anymore I'm impressed but now I was.
- NetMageSCW 10mo agoDepending on how important the GUI is to you, I would look into LINQPad for stuff that is scripting but too heavy.
- vednig 10mo agoPython's recent developments have been monumental, new versions now easily best the PyPy performance charts on M4 MacBook Air, idk if this has something to do with optimizations by Apple but coming from Linux I was surprised
- maximgeorge 10mo ago[dead]
- bboreham 10mo agoMatt Godbolt was saying recently that using tail-calls for an interpreter suits the branch predictor inside the cpu. Compared to a single big switch / computed jump.
- IshKebab 10mo agoI would have thought it actually helps the branch target predictor rather than the branch predictor. If you assume a simple predictor where the predicted target is just the last taken one then it's going to be wrong almost every time for a single switch. It will only be right for repeats of the exact same instruction. If you have a separate switch at the end of each instruction then it will be right any time an instruction is followed by the same instruction as last time, which can probably happen quite a lot for short loops.
- croemer 10mo ago2 typos in first sentence. Is this on purpose to make it obviously not-AI generated? "apology peice" and "tail caling"
- wk_end 10mo agoIf you want to make your writing appear non-AI generated, the easiest way is to write it yourself. No typos necessary. I’m sure with enough cajoling you can make the LLM spit out a technical blog post that isn’t discernibly slop - wanton emoji usage, clichés, self-aggrandizement, relentlessly chipper tone, short “punchy” paragraphs, an absence of depth, “it’s not just X—it’s a completely new Y” - but it must be at least a little tricky what with how often people don’t bother. [ChatGPT, insert a complaint about how people need to ram LLMs into every discussion no matter how irrelevant here.]
- eru 10mo ago> If you want to make your writing appear non-AI generated, the easiest way is to write it yourself. No typos necessary. You can ask the AI to make typos for you.
- kenjin4096 10mo agoWoops, thanks for noticing, fixed!
- wk_end 10mo agoSo…if the Python team finds tail calls useful, when are we going to see them in Python?
- dodomodo 10mo agoThey find them useful as a performance optimization, not as a design tool. This optimization is not relevant to Python code because it relies on the optimization passes the compiler makes.
- 01300415352 10mo ago[flagged]
- 01300415352 10mo ago[flagged]
- Shylathapa85 10mo ago[dead]
- hannaholovia 10mo ago[dead]
- johnbink 10mo ago[dead]
- s0a 10mo agostill python so it's only beating itself
- 3r7j6qzi9jvnve 10mo agoI now have to know why subparsers test got 60% slower... (0.3960 in the graph)
- malkia 10mo agoWow - clojure's recur in C/C++ - awesome!
- amai 10mo agoIs Python now faster than PHP? https://benchmarksgame-team.pages.debian.net/benchmarksgame/fastest/python3-php.html https://benchmarksgame-team.pages.debian.net/benchmarksgame/...