5 ms·
Quick overview of the design space: * PyPy JITs everything, so it can do _normal_ Python numerical code that is quite fast and regular Python code that is fast
by itamarst 4y ago
Quick overview of the design space:
* PyPy JITs everything, so it can do _normal_ Python numerical code that is quite fast and regular Python code that is fast. However, its interactions with libraries like NumPy add overhead, and it seems like it can't JIT code that interacts with NumPy in a useful way (AFAIK, would be happy to be proven wrong). So not useful for optimizing numeric functions that interact with libraries like NumPy.
* Plain old NumPy and friends. This is great... if the operation you want is already available as a "vectorized" API. "Vectorized" in this context does NOT mean SIMD, it's a Python-specific usage, see below.
* Numba: JIT compilation specifically focusing on interop with NumPy and similar libraries. Lets you write subset of Python but unlike NumPy you can use for loops and go fast.
* AOT compilation: Cython, Rust, C++, etc.. You have a longer feedback loop, but you have a full programming language, especially if you avoid Cython. OTOH Cython has nicer Python interop so for simple just-a-little-addon it can be easier to use if you don't already know Rust. You really shouldn't be writing new C++ in this day and age (but wrapping an existing library is useful). Like C++, Cython doesn't help with memory safety. Cython also suffers from two compilers, so debugging can be harder, especially if you use the C++ interop; if you are wrapping existing C++ library, I'd probably start with PyBind11 based on long-ago experience with Boost::Python.
Longer form:
* "Vectorization" in the context of Python: https://pythonspeed.com/articles/vectorization-python/ https://pythonspeed.com/articles/vectorization-python/
* PyPy and Numba as alternatives to vectorization: https://pythonspeed.com/articles/vectorization-python-alternatives/ https://pythonspeed.com/articles/vectorization-python-altern...
* Choosing a compiled language: https://pythonspeed.com/articles/rust-cython-python-extensions/ https://pythonspeed.com/articles/rust-cython-python-extensio...
* The performance overhead of AOT compiled libraries (less relevant if you're doing anything numeric): https://pythonspeed.com/articles/python-extension-performance/ https://pythonspeed.com/articles/python-extension-performanc...
* Numba intro: https://pythonspeed.com/articles/numba-faster-python/ https://pythonspeed.com/articles/numba-faster-python/
- hedgehog 4y agoGood overview. "Vectorized" is an old term that's been around since the early days of supercomputers and maybe before, not sure where it came from. Numba does a bunch of different things for code written to the Numpy API including CUDA acceleration. Certain machine learning frameworks like PyTorch and JAX also roughly follow the Numpy API because it is widely familiar and easy enough to work with. The kind of code that benefits from this kind of acceleration is hard to write yourself. A lot of workloads lean on linear algebra operations that are conceptually simple but complicated to implement with good performance, thus why all of this tooling isn't just a couple thousand lines of C. Good overview of matmul on CPU: https://gist.github.com/nadavrot/5b35d44e8ba3dd718e595e40184d03f0 https://gist.github.com/nadavrot/5b35d44e8ba3dd718e595e40184...
- college_physics 4y agoCray supercomputers used to have special "vector" units that would perform operations on multiple scalars (eg 128 doubles) in parallel. A bit like gpu units. Any algorithm that could be cast in a form that benefited from this type of parallelism would be called vectorizable. Linear algebra obviously fits perfectly (but, depending on the problem, you might0 need to juggle the vector dimensions). Vectorizing code was fairly straightforward using the latter versions of fortran. It was all quite sweet and productive but could not provide the required hpc scaling so was eventually abandoned in favor of massively parallel designs.
- hedgehog 4y agoI think this is backwards, the early Crays and maybe other machines used essentially serial + aggressively pipelined CPUs while modern CPUs do actually execute operations in parallel via multiple ALUs but are from a programmer point of view kind of similar (fixed size vector).
- college_physics 4y agoWhat is a vector unit was is simd etc can be confusing but at least some wikipedians are sure the early Crays had true vector processing (which may not be parallel processing according to your definition) https://en.m.wikipedia.org/wiki/Vector_processor https://en.m.wikipedia.org/wiki/Vector_processor
- hedgehog 4y agoYes, in the Cray design a vector instruction processes one item per clock so total latency increases with number of elements. Only parallel in the sense of pipelining being a form of parallelism. Something like an AVX-equipped Intel CPU processes all elements in parallel to deliver a result in an essentially constant number of cycles. Edit: There's a period write-up of the general Cray 1 design here: https://inst.eecs.berkeley.edu/~n252/sp07/Papers/Cray.pdf https://inst.eecs.berkeley.edu/~n252/sp07/Papers/Cray.pdf