5 ms·
Rather than go to all this hassle, I think it’s much easier to use Cython, or even to write C and just wrap it with Cython/Ctypes. Performance is so much better
by physicsguy 2y ago
Rather than go to all this hassle, I think it’s much easier to use Cython, or even to write C and just wrap it with Cython/Ctypes. Performance is so much better.
- wdroz 2y agoI agree, you can also use Rust with PyO3[0] and Maturin[1]. It's far easier than what the author is doing is the article. [0] -- https://github.com/PyO3/pyo3 https://github.com/PyO3/pyo3 [1] -- https://github.com/PyO3/maturin https://github.com/PyO3/maturin
- Austizzle 2y agoI've done this and it's a really lovely way to work once the rust stuff clicks!
- KeplerBoy 2y agoyes, generating Python bytecode seems to be a horrible idea, and 2x improvements are not great coming from Python. Apart from Cython and other C/C++ bindings one might also want to look at JIT compilers like Numba (or Jax and Pytorch for number crunching).
- willseth 2y agoThe improvement was 2x across the entire application, i.e. this hotspot was 50% of the overall runtime, and his bytecode optimization reduced that to an insignificant %. The improvement to the code he actually optimized was probably at least 2 orders of magnitude.
- mmoskal 2y agoTypical speed difference between Python and native is around two orders of magnitude (of course for code that actually does computation in Python, not just call numpy etc).
- fire_lake 2y agoIt can actually be way more than that for some applications - I found 1000x for some algorithm heavy use cases. I don’t find Python ergonomic enough to justify the performance penalty.
- willseth 2y agoIme 2 orders of magnitude is closer to best case, assuming you have already made reasonable Python-based optimizations. But this problem has two components, function call overhead and conditionals. Moving the original logic into a compiled extension would eliminate a lot of interpreter overhead for function calls, but it would not eliminate the branchy code. His codegen eliminated both.
- IshKebab 2y agoNah 2 orders of magnitude is very typical when going from Python to a "fast" language like C++ or Rust.
- willseth 2y agoIt's pointless to argue what's "typical" when actual results are incredibly varied and specific to each operation. Also you didn't read the rest of my comment.
- IshKebab 2y agoThey're varied. But 2 orders of magnitude is typical. Sometimes it's 1, sometimes 3. Rarely 0 or 4. I did read the rest of your comment; I was just not responding to that part.
- pjmlp 2y agoOr use a language with JIT/AOT in the box, REPL support, some of which even have IDEs all the way back to 1990's.
- hyperpape 2y agoWhile it's probably less often used, runtime code generation is a potential technique for optimizing performance in Java. Elsewhere iirc, V8, which is written in C++, can use runtime code generation for regexes.
- stefanos82 2y agoOr you could start with mypyc [1] first from mypy and if the results do not please you, then you can use Cython; some folks don't want to learn yet another programming syntax, even though it's Pythonic enough to learn it in a rather short period of time, but still... https://github.com/python/mypy https://github.com/python/mypy
- kevin_thibedeau 2y agoCython can speed up plain Python code. It bypasses the interpreter for everything that is PyObject-based.
- olejorgenb 2y ago(https://github.com/mypyc/mypyc https://github.com/mypyc/mypyc) That's cool! > Mypyc compiles Python modules to C extensions. It uses standard Python type hints to generate fast code. Mypyc uses mypy to perform type checking and type inference. > > Mypyc can compile anything from one module to an entire codebase. The mypy project has been using mypyc to compile mypy since 2019, giving it a 4x performance boost over regular Python.
- willseth 2y agoThe problem isn't simply that repeated function calls are expensive, it's that the comparator logic is very branchy in order to support the different comparator options for their queries, even though any given query will only perform the same comparator for every value. The operation being too dynamic would exist regardless. He jumps straight from the problem to codegen-based solutions, but I wonder if a simpler loop specialization would yield a big enough chunk of the performance he got. If you hoist the comparator handling higher up into a single branch per query, then have a separate loop for each, you avoid unnecessary branching in your hot loop. You'd have to duplicate logic for every comparator, which is a little ugly, but there were only a few comparators, and that seems a lot more palatable if the alternative is an elaborate codegen solution.