12 ms·
Speed up your Python using Rust
- pmoriarty 9y agoWhat's the advantage of doing this over using cython or pypy?
- sametmax 9y agoMore perfs since you are going lower level. You are not limited by cython types or pypy git warming up. You can use the full range of c libs AND rust libs. You benefit from the cargo build tools.
- oakridge 9y agoI'm not sure how different this is with Cython since you can also write C/C++ code with it, and hence use any C/C++ libraries. I say that the benefit offered by the example is more of bringing the Rust ecosystem to Python than solving performance issues.
- sametmax 9y agoWell, with cython, either you write python compiled to C, and you won't match rust perf nor types, or you write c/c++ and bind it to python, and you won't match rust safetiness.
- rjeli 9y agowhy would rust have better performance than cython? cython is as fast as C
- coldtea 9y agoCython just transpiles to C. It's not "as fast as C", it's as fast as the runtime support structures it uses, and the communication with Python allows.
- rjeli 9y agoyes, exactly as I said.
- ogrisel 9y agoYou can write very efficient Cython code but it's true that in this case, you tend to adopt a lower level code style that is very close to C/C++. Basically, you need to think about the C/C++ code that will be generated by Cython. C/C++ compilers might be able to generate more optimized native code than what rutsc does though. Actually, this is a question: how good is rustc with numerical / math intensive code? For instance, does it implements loop unrolling and SIMD vectorization?
- heavenlyblue 9y agoMost of the time when a developer needs loop unrolling - numpy will work best anyway. Why does everyone always start mentioning this fact when performance is mentioned? For example, in my case I always need high-performance code to work with strings loaded from loads of CSV files. That includes: merging strings, matching them, comparing them. Loop unrolling/SIMD would not really help here, while an ability to write safe, checked code fast - would. On the other hand I do need the pythonic dynamics, so that's what I stick to.
- toyg 9y agoTotal layman here, but I thought one would use Rust because it’s easier to write “safe” code with it than C/C++, while maintaining an equivalent low-level speed advantage.
- baq 9y agoyou should also quote "easier" ;) it's certainly easier once you sacrifice a goat to the compiler to stop insulting you and your mother. disclaimer: i love rust
- sametmax 9y agoWell it's better than your client insulting you and your mother because your software killed their goat by mistake.
- danilocesar 9y agolol. It's sad that I can only upvote this once :)
- fulafel 9y agoCython is unsafe.
- fermigier 9y agoIt's much easier to learn Cython once you know Python. That's the biggest selling point.
- the_mitsuhiko 9y agoBut the tooling is terrible in comparison. I find rust sigificantly easier as a python developer than cython. There is so much more rust ecosystem to take advantage of.
- tdbgamer 9y agoTo be fair to Cython, they have access to the entire C++ stdlib, so that's a fairly good amount of tooling. The main thing is lacks is good documentation and memory safety.
- the_mitsuhiko 9y agoAnd C++ has absolutely no package distribution system at this point.
- tdbgamer 9y agoCython deceptively similar to Python with a lot of the pitfalls of C baked in. I personally found their documentation lacking and had to read through tons of Cython projects to discover how it behaved in many scenarios. I've had a much better experience with Rust documentation in comparison.
- conn01 9y agoRust is disgusting language with too much commercial exploitation
- 9y ago
- hasenj 9y agoSpeed up your python by not using python at all. Learn a language that compiles to native binaries and be better off for it. For example: Go, Swift, D.
- sdfjkl 9y agoNim probably makes more sense if you're coming from Python. Oh, and Python is perfectly fine for most things. If you like it, sticking with it and just speeding up the rare thing that needs that performance makes much more sense than switching language, especially if you already are well into a project.
- rochacbruno 9y agoYeah, that is exactly what I said in the article! Speed up your Python using Rust; Only for those rare cases when Python is detected as the bottleneck. And of course, it can be C/C++ or Nim or Go etc...
- tatterdemalion 9y agoOr Rust. ;-)
- vog 9y agoI like the article, but the following advice confused me, especially since this comes from RedHat i.e. Linux people: > Having Rust installed (recommended way is https://www.rustup.rs/ https://www.rustup.rs/). This essentially recommends unconditionally using the "curl | sh" anti-pattern. Shouldn't they recommend instead e.g. "apt-get install rustc" for Debian users? Since this doesn't make use of too recent Rust features, using Rust 1.14 of Debian/Stable should be fine, shouldn't it? Same of Fedora, etc.
- cesarb 9y ago> using Rust 1.14 of Debian/Stable should be fine, shouldn't it? Same of Fedora, etc. That is a RedHat developer blog. Is rust available on RHEL already? Looking at CentOS (which should have nearly the same packages), it doesn't appear to be available yet.
- steveklabnik 9y agoThis stuff is in preview https://developers.redhat.com/blog/2017/10/04/red-hat-adds-go-clangllvm-rust-compiler-toolsets-updates-gcc/ https://developers.redhat.com/blog/2017/10/04/red-hat-adds-g...
- DC-3 9y agoRustup is very useful for managing toolchains. For someone new to Rust, it's probably best to be familiar with it from the outset.
- simias 9y agoIf you look at the way the rustup-init.sh script is written it's safe to be used with this "anti pattern". I see your objection though but unfortunately this ship has sailed, you might as well complain about websites that don't work without Javascript... The advantage of this method is that it will work on any linux distro (and even BSD, Darwin and mingw) and you'll get the latest stable version. I don't see the advantage of using potentially outdated OS packages for installing a compiler, it's not like it's a dependency for other packages. It also makes it easy to manage the various components of the toolchain, for instance if you later want to crosscompile for an other target, use the nightly version etc...
- Matumio 9y agoFor comparison, I just implemented the same as C SWIG extension[1]. It's about 10% faster, but it's cheating by comparing bytes instead of utf-8 encoded characters. The more interesting part to me is the comparison of the amount of boilerplate code required. https://github.com/martinxyz/rust-python-example/commit/f8e36ab5f9c https://github.com/martinxyz/rust-python-example/commit/f8e3...
- emj 9y agopytest.benchmark really needs to default to a smaller width of it's stats, those stats are really just meant to be used in a terminal..
- nabla9 9y agoI'm not familiar with Rust libraries, but I would guess it's just counting code points and not characters, so strictly speaking both are cheating. I would love to see people showing how to do simple string processing, like counting characters in proper grapheme cluster level in their favorite programming language.
- steveklabnik 9y agofor (c1, c2) in val.chars().zip(val.chars().skip(1)) chars() iterates by unicode scalar values. It'd be bytes() for bytes. If you wanted to do it by grapheme clusters, you'd add https://crates.io/crates/unicode-segmentation https://crates.io/crates/unicode-segmentation to your Cargo.toml, add the relevant imports you see on that page to your code, and change the above line to for (c1, c2) in UnicodeSegmentation::graphemes(val, true).zip(UnicodeSegmentation::graphemes(val, true).skip(1)) ... possibly splitting that up into variables becuase dang, that's a long line. Then, you're getting &strs instead of chars for the iteration, but I think the body still says the same, as == checks by value.
- radarsat1 9y agoOne thing though that gets very complicated about using SWIG is ownership semantics. With anything more complicated than passing scalar values, it is very easy to introduce a memory leak or double-free if you don't get the flags right. I wonder if Rust types naturally allow a much better inference of ownership semantics across the language boundary?
- emj 9y agoAbout as fast as numpy.. More tools to create fast code is always great, but the tooling for Rust/C in Python needs to be easier, I just can't be bothered most of the time. This in numpy gets a better relative boost on my machine YMMV. import numpy def count_double_chars_np(val): ng=np.fromstring(val,dtype=np.byte) return np.sum(ng[:-1]==ng[1:]) def test_np(benchmark): benchmark(count_double_chars_np, val)
- dr_zoidberg 9y agoGood numpy implementation of the algorithm. If, for whatever reason, numpy isn't available, you can also pull it with a good comprehension: def count_doubles2(val): return sum(1 for c1, c2 in zip(val, val[1:]) if c1 == c2) Which will also allow you to avoid a function call entirely, if it was useful in some way: In [56]: %timeit count_doubles(val) 198 ms ± 13.8 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) In [57]: %timeit count_doubles2(val) 189 ms ± 21.1 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) In [58]: %timeit sum(1 for c1, c2 in zip(val, val[1:]) if c1 == c2) 135 ms ± 3.86 ms per loop (mean ± std. dev. of 7 runs, 10 loops each) In [59]: %timeit count_double_chars_np(val) 6.95 ms ± 782 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) (Numpy still beats it, for long strings).
- rochacbruno 9y agoHi, can you send a Pull Request including your numpy implementation? https://github.com/rochacbruno/rust-python-example https://github.com/rochacbruno/rust-python-example I would like to add it there just for the record and then I will update the article.
- emj 9y agoThanks would be nice, you should be able to just copy and paste that oneliner, but I'm not sure you blog post is better for it. The idea that Rust is easier to include in Python is important enough, and Numpy is a bit of an edge case imho. Ideas are always CC-zero
- fnl 9y agoThank you for the very nice, educative article, Bruno! If performance comparison of counting character pairs really were the issue here, in addition to the already suggested numpy approach, an implementation I'd dare wager to be as competitive is re2, e.g. [1], a drop-in replacement for the standard re package. But I want to point out that I think all this performance comparison of this trivial character counting distracts from the core idea here: You'd use a low-level implementation in Rust (or C/C++/Cython, for that matter) when such "nifty tricks" are not available, after all. So again thanks for the article, and do think if you really want this performance issues to degrade the article to a only marginally relevant performance "showdown". https://pypi.python.org/pypi/re2/ https://pypi.python.org/pypi/re2/
- b0rsuk 9y ago> Rust is a language that, because it has no runtime, can be used to integrate with any runtime; you can write a native extension in Rust that is called by a program node.js, or by a python program, or by a program in ruby, lua etc. and, however, you can script a program in Rust using these languages. — “Elias Gabriel Amaral da Silva” Can someone explain why is "having a runtime" problematic for writing extensions and calling them from Python ? From what I gather Go does have a runtime, so implicitly it should be suboptimal for calling from Python. Yet since 2015 (Go 1.5) can be called directly from Python. I'm a Python programmer looking to expand my tool belt. I'm wondering of relative pros and cons of Rust and Go. I have only written small toy programs in C and other compiled languages. Is Go better suited to completely rewriting software rather than using it for extensions ? Why ? I would appreciate a benchmark with a Go extension, too.
- sigzero 9y agoMaybe? https://matthias-endler.de/2017/go-vs-rust/ https://matthias-endler.de/2017/go-vs-rust/
- forkerenok 9y agoIf I understand the matter correctly, FFI-ing with a language that has runtime has more friction in extra overhead of initializing runtime, i.e. more state management in your app. Not that it is not doable. EDIT: maybe indeed someone more knowledgeable will explain it or point to good condensed reads.
- jerf 9y ago"Can someone explain why is "having a runtime" problematic for writing extensions and calling them from Python ?" Perhaps instead of saying "having a runtime" it would be better to examine the situation in terms of what the code assumes. Python assumes that it has the Python GC running on its code, that everything is a PyObject of one sort or another, that it has a Global Interpreter Lock that if taken will prevent anything from modifying anything it thinks it owns, and so on. Go assumes that it has the Go GC running (despite both "having GC", there's enough differences that it must be specified as a difference), that its objects are laid out in certain manners such that most field references are compiled down to static offsets rather than dynamic lookups, that it can run its core event loop and dispatch out work to its internal goroutines without asking anyone else, etc. You could go on for quite a while; I don't intend those as complete lists. I just want to convey the flavor of conceptualizing the runtime in terms of assumptions that the code running in that runtime can make. Once you look at it this way, it should be more clear why trying to jam two runtimes into one OS process gets to be tricky. I use the word "jam" quite carefully, because it always feels that way to me. The more differences between the assumptions of the two runtimes, the more translation the code is going to need. For instance, Python to anything else is going to involve unwrapping the data from the internal PyObject wrappers, and wrapping anything coming back from somewhere else back into PyObjects. Threading models have to be matched up. Memory layout has to be harmonized. Memory generally has to be kept strictly separated, because the two runtimes both expect to be able to manage memory, so you can't hand memory allocated by one of them to the other, which further implies that you're almost certainly copying everything across the boundary. Etc. etc. I'd also separate out the way there can be differences in the affordances of the languages. For instance, Python doesn't have what Rust or Go would call "arrays". Rust and Go are fine with getting arrays of pointers, but the languages afford the use of memory-contiguous arrays without pointers, so especially if you're integrating with a third-party library, you have no choice but for some layer somewhere along the way to convert Python lists into the correct sort of array. The runtimes technically don't force this, but the structure of the libraries and code afforded by the other languages do. By contrast, if you were integrating with lisp, you might find many points where you need to turn things into singly-linked lists, again, not because Lisp can't handle arrays, but because you're likely to encounter pre-existing Lisp code that expects Lisp cons lists. As another example, despite the fact Go and C generally see eye-to-eye on how to layout structs, the C support from Go is still extremely expensive due to the need to convert from how Go sees the concurrency world to how C sees the world. C, contrary to popular belief, actually does have a runtime, and that runtime tends to assume it has very deep control of the OS process it is running in. Go has to do a lot of work to isolate the running C code in an environment it is comfortable with, where it won't be pre-empted by the green thread code (on account of the fact that it can't be, C doesn't support that). There's also some tricksy code you may need to write to harmonize C's memory-management-via-malloc model with Go's "lifetimes determined via the GC" model. (If you listen carefully, you can hear the Go runtime go "klunk" every time it runs cgo code.) Rust has a runtime too, but unlike a lot of languages, it has the ability to shut it off. You lose some services and capabilities, but on the upside, you significantly reduce the number of assumptions the Rust code is making, making it easier to integrate with other runtimes. (I say reduce because technically, it still doesn't make it to zero if you are precise enough in your thinking, but I'd expect that of all the current "cool" languages, Rust with the runtime off probably makes fewer assumptions than anything else.) That said, I'm not sure if this code is working in that mode. I see the rust code doesn't directly turn off the runtime, but I don't know what that "#[macro_use] extern crate cpython;" line fully expands to. It's possible that the full Rust runtime is still in play, which looks enough like C anyhow (by explicit design of the Rust team) that Python's existing C integration can just be reused. Either way Rust is still making many fewer assumptions that Go's relatively heavyweight (in terms of assumptions moreso than resources) runtime.
- js2 9y agoSee also “Fixing Python Performance With Rust” previously discussed here: https://news.ycombinator.com/item?id=12748020 https://news.ycombinator.com/item?id=12748020 And “Evolving Our Rust With Milksnake”: https://news.ycombinator.com/item?id=15697570 https://news.ycombinator.com/item?id=15697570 Both from Armin Ronacher at Sentry. About Milksnake: Milksnake helps you compile and ship shared libraries that do not link against libpython either directly or indirectly. This means it generates a very specific type of Python wheel. Since the extension modules do not link against libpython they are completely Python version or implementation independent. The same wheel works for Python 2.7, 3.6 or PyPy. As such if you use milksnake you only need to build one wheel per platform and CPU architecture.
- rochacbruno 9y agoYeah, Milksnake is mentioned in the article :)
- js2 9y agoDoh, completely missed that reading the article on my phone.
- onnnon 9y agoIf anyone is looking for something like this for Ruby, check out Helix: https://usehelix.com https://usehelix.com
- jD91mZM2 9y agoAwesome, thanks! It feels like a lot of people dislike Ruby, but I'm glad some people still like it. I think it's a good Python alternative, and has syntax that reminds you of Rust.
- Dowwie 9y agoI never got around to talking about it, but as part of my "month of Rust", I ported permission-based authorization logic from Python to Rust and then ran performance benchmarks of the Rust implementation and a pypy-compiled version. The pypy-compiled python ran slightly faster. I've been told not to expect similar results in other implementations. These findings cannot be used to draw any conclusions about pypy. My rust project: https://github.com/YosaiProject/yosai_libauthz https://github.com/YosaiProject/yosai_libauthz
- staticassertion 9y agoThat's pretty interesting - JIT's can do a lot of great optimizations with runtime information, but I'm still surprised to hear that Pypy was faster. It would be cool to see the benchmarks, methodology, and Python code. Personally, Pypy has never been an option due to the nature of the codebases I work on - or at least it wasn't. I was using pandas, numpy, scipy etc and I don't think it was compatible.
- k__ 9y agoI don't understand all the implications, but I often hear from JIT language people that their language could be as fast as a AOT language if it was used correctly. What are the use-cases that lend itself to be faster in Rust, than JS/Ruby/Phython?
- Dowwie 9y agoI would like to know this as well. This code is one instance of it.
- brightball 9y agoFwiw, Ruby does get much more performant with JRuby but nobody cares that much because the benefits are mostly lost with Rails.
- pc2g4d 9y agoVery interesting. I glanced at the code to see if there were obvious performance issues. All I noticed was a triple `map` invocation in https://github.com/YosaiProject/yosai_libauthz/blob/master/src/c_abi.rs https://github.com/YosaiProject/yosai_libauthz/blob/master/s... Not sure what the compiler does with that, but I'd expect that it means you're running through those elements three times. I imagine it would be better to reduce this to one `map` call.
- pbreit 9y agoShouldn't Python, et al have more "native" ways to achieve these sorts of performance improvements?
- rochacbruno 9y agoMaybe because of the existence of `Numpy` and `Cython` and `PyPy` + all other `FFI` possibilities it is not on Python Roadmap.
- j_s 9y agohttps://news.ycombinator.com/item?id=14588333 https://news.ycombinator.com/item?id=14588333 (beautifulsoup/lxml upgrade) >Python: interactive glue language between high performance C libraries. Appreciate this walkthrough for Rust!
- rochacbruno 9y agoHow does it compare with https://github.com/servo/html5ever https://github.com/servo/html5ever (someone with free time do run some benchmarks)
- j_s 9y agoThat's a great question, and fits nicely in the context of the current discussion. I think the primary claim to fame for this C-based https://github.com/kovidgoyal/html5-parser https://github.com/kovidgoyal/html5-parser is serving as a drop-in performance boost for lxml (at the API level; it parses invalid HTML differently/more consistently). I too would be interested in a performance comparison to help decide which project makes more sense for new projects. The existing Python layer in html5-parser might give it a leg up if the language of choice is Python - is there a similar project for the Rust-based html5ever?
- rochacbruno 9y agoNow we got Numba, Cython and Numpy results for comparison https://github.com/rochacbruno/rust-python-example#new-results https://github.com/rochacbruno/rust-python-example#new-resul...