6 ms·
I've been playing with PyO3 for prototyping, and wrapped some Rust code to see if it's faster than Python. The experience was very much like using Boost Python
by itamarst 6y ago
I've been playing with PyO3 for prototyping, and wrapped some Rust code to see if it's faster than Python. The experience was very much like using Boost Python (whcih these days has alternative with https://github.com/pybind/pybind11 https://github.com/pybind/pybind11). It's _really_ easy to wrap code for Python, and it has nice APIs to ensure GIL is held. Being Rust, I'm much more confident I won't suffer from memory unsafety issues which my C++ at the time did.
Now I'm starting to use it as part of the Python memory profiler I'm working on (https://pythonspeed.com/fil https://pythonspeed.com/fil), in this case to call in to the low-level Python C API which PyO3 includes bindings for in addition to its high-level API. This kind of usage is more like writing C, except with the benefit of having high-level APIs (for GIL holding, but also object conversion) available when I need it.
So basically you get safe, high-level, easy-to-use APIs, with fallback to low-level unsafe APIs if you need them.
Highly recommend trying it out.
- shirakawasuna 6y agoSounds great! Would so much rather drop into Rust than C or C++.
- brundolf 6y agoWhat's the data-conversion overhead look like at the boundary? Which data structures can be passed back and forth without a full clone, etc?
- itamarst 6y agoThere's definitely a conversion cost. For strings, Python apparently caches the UTF-8 encoded string, so if you _repeatedly_ transfer it to Rust I suspect (but haven't checked) that the cost is much lower. In general I suspect it's the usual "NumPy arrays are fast, everything else you better be getting a sufficiently large boost from the low-level code to justify conversion". For the thing I prototyped in Rust, it was wrapping the `ahocorasick` crate which was in fact faster than `pyahocorasick` which is written in C or Cython or something. Both have similar conversion costs, probably, so it came down to "for lots of data the Rust version was faster".
- burntsushi 6y agoBe sure to use auto configuration to get it to go even faster, depending on your use case: https://docs.rs/aho-corasick/0.7.15/aho_corasick/struct.AhoCorasick.html#method.new_auto_configured https://docs.rs/aho-corasick/0.7.15/aho_corasick/struct.AhoC... Or just be sure to enable the DFA option if you can afford it. It looks like the Python library is just the standard NFA algorithm.
- itamarst 6y agoYeah, I was using DFA. Next step is trying alternative approach, but if that alternative doesn't work I'm going to see about wrapping your package for Python. Thanks for all your work on it!
- burntsushi 6y agoNice! Reach out if there are any problems or if you need something exposed in the API. Looking at the pyahocorasick issue tracker, there are a number of features/bugs that your wrapper package would resolve. :)
- liuliu 6y agoNumPy also support conversions without copying. One thing I haven't found good way to bridge between Python is the pandas.DataFrame, it seems to be quite Python focused object and iterating through DataFrame is particularly slow.
- itamarst 6y agoInternally Pandas often uses NumPy arrays, especially for numeric data, so might be able to pass things that way in some cases? E.g. `df["column_name"].values` will you get you a NumPy array.
- JPKab 6y agoWas just checking out your fil project. It looks really useful, and I dig the jupyter kernel as well.
- itamarst 6y agoThank you! If you have any questions/problems/ideas, please reach out via GitHub or email (itamar@pythonspeed.com).