49 ms·
Using Rust in non-Rust servers to improve performance
- bebna 2y agoFor me a "Non-Rust Server" would be something like a PHP webhoster. If I can run my own node instance, I can possible run everything I want.
- bluejekyll 2y agoThe article links to two PHP and Rust integration strategies, WASM[1] or native[2]. [1] https://github.com/wasmerio/wasmer-php https://github.com/wasmerio/wasmer-php [2] https://github.com/davidcole1340/ext-php-rs https://github.com/davidcole1340/ext-php-rs
- eandre 2y agoEncore.ts is doing something similar for TypeScript backend frameworks, by moving most of the request/response lifecycle into Async Rust: https://encore.dev/blog/event-loops https://encore.dev/blog/event-loops Disclaimer: I'm one of the maintainers
- internetter 2y agoWhat's your response to this? https://github.com/encoredev/ts-benchmarks/issues/2 https://github.com/encoredev/ts-benchmarks/issues/2
- uncomplexity 2y agonot gp bot first time seeing this encore ts. i've been a user of uwebsockets.js, uwebsockets is used underneath by bun. i hope encore does benchmark compared to encore, uwsjs, bun, and fastify. express is just so damn slow. https://github.com/uNetworking/uWebSockets.js https://github.com/uNetworking/uWebSockets.js
- eandre 2y agoWe've published benchmarks against most of these already, see https://github.com/encoredev/ts-benchmarks https://github.com/encoredev/ts-benchmarks
- eandre 2y agoI've published proper instructions for benchmarking Encore.ts now: https://github.com/encoredev/ts-benchmarks/blob/main/README.md https://github.com/encoredev/ts-benchmarks/blob/main/README..... Thanks!
- isodev 2y agoThis is a really cool comparison, thank you for sharing! Beyond performance, Rust also brings a high level of portability and these examples show just how versatile a pice of code can be. Even beyond the server, running this on iOS or Android is also straightforward. Rust is definitely a happy path.
- jvanderbot 2y agoRust deployment is a happy path, with few caveats. Writing is sometimes less happy than it might otherwise be, but that's the tradeoff. My favorite thing about Rust, however, is Rust dependency management. Cargo is a dream, coming from C++ land.
- krick 2y agoEverything is a dream, when coming from C++ land. I'm still incredibly salty about how packages are managed in Rust, compared to golang or even PHP (composer). crates.io looks fine today, because Rust is still relatively unpopular, but 1 common namespace for all packages encourages name squatting, so in some years it will be a dumpster worse than pypi, I guarantee you that. Doing that in a brand-new package manager was incredibly stupid. It really came late to the market, only golang's modules are newer IIRC (which are really great). Yet it repeats all the same old mistakes.
- Imustaskforhelp 2y agoIn my opinion , I like golang's way better because then you have to be thoughtful about your dependencies and it also prevents any drama (like rust foundation cargo drama) (ahem) (if you are having a language that is so polarizing , it would be hard to find a job in that ) I truly like rust as a performance language but I would rather like real tangible results (admittedly slow is okay) than imagination within the rust / performance land. I don't want to learn rust to feel like I am doing something "good" / "learning" where I can learn golang at a way way faster rate and do the stuff that I like for which I am learning programming. Also just because you haven't learned rust doesn't make you inferior to anybody. You should learn because you want to think differently , try different things. Not for performance. Performance is fickle minded. Like I was seeing a native benchmark of rust and zig (rust won) and then I was seeing benchmark of deno and bun (bun won) (bun is written in zig and deno in bun) The reason I suppose is that deno doesn't use actix and non actix servers are rather slower than even zig. It's weird .
- bhelx 2y agoIf you have a Java library, take a look at Chicory: https://github.com/dylibso/chicory https://github.com/dylibso/chicory It runs on any JVM and has a couple flavors of "ahead-of-time" bytecode compilation.
- deleted 2y ago[deleted]
- bluejekyll 2y agoThis is great to see. I had my own effort around this that I could never quite get done. I didn’t notice this on the front page, what JVM versions is this compatible with?
- evacchi 2y agoJava 11+ :)
- bluejekyll 2y agoPerfect!
- xyst 2y agoIn my opinion, the significant drop in memory footprint is truly underrated (13 MB vs 1300 MB). If everybody cared about optimizing for efficiency and performance, the cost of computing wouldn’t be so burdensome. Even self-hosting on an rpi becomes viable.
- leeoniya 2y agofwiw, Bun/webkit is much better in mem use if your code is written in a way that avoids creating new strings. it won't be a 100x improvement, but 5x is attainable.
- dangsux 2y ago[dead]
- jchw 2y agoIt's a little more nuanced than that of course, a big reason why the memory usage is so high is because Node.JS needs more of it to take advantage of a large multicore machine for compute-intensive tasks. > Regarding the abnormally high memory usage, it's because I'm running Node.js in "cluster mode", which spawns 12 processes for each of the 12 CPU cores on my test machine, and each process is a standalone Node.js instance which is why it takes up 1300+ MB of memory even though we have a very simple server. JS is single-threaded so this is what we have to do if we want a Node.js server to make full use of a multi-core CPU. On a Raspberry Pi you would certainly not need so many workers even if you did care about peak throughput, I don't think any of them have >4 CPU threads. In practice I do run Node.JS and JVM-based servers on Raspberry Pi (although not Node.JS software that I personally have written.) The bigger challenge to a decentralized Internet where everyone self-hosts everything is, well, everything else. Being able to manage servers is awesome. Actually managing servers is less glorious, though: - Keeping up with the constant race of security patching. - Managing hardware. Which, sometimes, fails. - Setting up and testing backup solutions. Which can be expensive. - Observability and alerting; You probably want some monitoring so that the first time you find out your drives are dying isn't months after SMART would've warned you. Likewise, you probably don't want to find out you have been compromised after your ISP warns you about abuse months into helping carry out criminal operations. - Availability. If your home internet or power goes out, self-hosting makes it a bigger issue than it normally would be. I love the idea of a world where everyone runs their own systems at home, but this is by far the worst consequence. Imagine if all of your e-mails bounced while the power was out. Some of these problems are actually somewhat tractable to improve on but the Internet and computers in general marched on in a different more centralized direction. At this point I think being able to write self-hostable servers that are efficient and fast is actually not the major problem with self-hosting. I still think people should strive to make more efficient servers of course, because some of us are going to self-host anyways, and Raspberry Pis run longer on battery than large rack servers do. If Rust is the language people choose to do that, I'm perfectly content with that. However, it's worth noting that it doesn't have to be the only one. I'd be just as happy with efficient servers in Zig or Go. Or Node.JS/alternative JS-based runtimes, which can certainly do a fine job too, especially when the compute-intensive tasks are not inside of the event loop.
- dyzdyz010 2y agoMake Rustler great again!
- jchw 2y agoHaha, I was flabbergasted to see the results of the subprocess approach, incredible. I'm guessing the memory usage being lower for that approach (versus later ones) is because a lot of the heavy lifting is being done in the subprocess which then gets entirely freed once the request is over. Neat. I have a couple of things I'm wondering about though: - Node.js is pretty good at IO-bound workloads, but I wonder if this holds up as well when comparing e.g. Go or PHP. I have run into embarrassing situations where my RiiR adventure ended with less performance against even PHP, which makes some sense: PHP has tons of relatively fast C modules for doing some heavy lifting like image processing, so it's not quite so clear-cut. - The "caveman" approach is a nice one just to show off that it still works, but it obviously has a lot of overhead just because of all of the forking and whatnot. You can do a lot better by not spawning a new process each time. Even a rudimentary approach like having requests and responses stream synchronously and spawning N workers would probably work pretty well. For computationally expensive stuff, this might be a worthwhile approach because it is so relatively simple compared to approaches that reach for native code binding.
- tln 2y agoThe native code binding was impressively simple! 7 lines of rust, 1 small JS change. It looks like napi-rs supports Buffer so that JS change could be easily eliminated too.
- jchw 2y agoI've used napi-rs a bit ago, it's pretty awesome. That said though, the main issue is that the Rust bindings story is not always that nice. It really depends. Internally, Node modules have quite a lot of complexity, and when you try to do more interesting things you could wind up facing some of the complexity of how it is implemented.
- tialaramex 2y agoCaveman approach has several nice features - I think I'd be tempted even if it didn't have better performance.
- deleted 2y ago[deleted]
- echelon 2y agoRust is simply amazing to do web backend development in. It's the biggest secret in the world right now. It's why people are writing so many different web frameworks and utilities - it's popular, practical, and growing fast. Writing Rust for web (Actix, Axum) is no different than writing Go, Jetty, Flask, etc. in terms of developer productivity. It's super easy to write server code in Rust. Unlike writing Python HTTP backends, the Rust code is so much more defect free. I've absorbed 10,000+ qps on a couple of cheap tiny VPS instances. My server bill is practically non-existent and I'm serving up crazy volumes without effort.
- manfre 2y ago> I've absorbed 10,000+ qps on a couple of cheap tiny VPS instances. This metric doesn't convey any meaningful information. Performance metrics need context of the type of work completed and server resources used.
- boredumb 2y agoI've been experimenting with using Tide, sqlx and askama and after getting comfortable, it's even more ergonomic for me than using golang and it's template/sql librarys. Having compile time checks on SQL and templates in and of itself is a reason to migrate. I think people have a lot of issues with the life time scoping but for most applications it simply isn't something you are explicitly dealing with every day in the way that rust is often displayed/feared (and once you fully wrap your head around what it's doing it's as simple as most other language aspects).
- JamesSwift 2y ago> Writing Rust for web (Actix, Axum) is no different than writing Go, Jetty, Flask, etc. in terms of developer productivity. It's super easy to write server code in Rust. I would definitely disagree with this after building a micro service (url shortener) in rust. Rust requires you to rethink your design in unique ways, so that you generally cant do things in the 'dumbest way possible' as your v1. I found myself really having to rework my design-brain to fit rusts model to please the compiler. Maybe once that relearning has occurred you can move faster, but it definitely took a lot longer to write an extremely simple service than I would have liked. And scaling that to a full api application would likely be even slower. Caveat that this was years ago right when actix 2 was coming out I believe, so the framework was in a high amount of flux in addition to needing to get my head around rust itself.
- voiper1 2y agoWow, that's an incredible writeup. Super surprised that shelling out was nearly as good any any other method. Why is the average bytes smaller? Shouldn't it be the same size file? And if not, it's a different alorithm so not necessarily better?
- pixelesque 2y ago> Why is the average bytes smaller? Shouldn't it be the same size file? The content being encoded in the PNG was different ("https://www.reddit.com/r/rustjerk/top/?t=all https://www.reddit.com/r/rustjerk/top/?t=all" for the first, "https://youtu.be/cE0wfjsybIQ?t=74 https://youtu.be/cE0wfjsybIQ?t=74" for the second example - not sure whether the benchmark used different things?), so I'd expect the PNG buffer pixels to be different between those two images and thus the compressed image size to be a bit different, even if the compression levels of DEFLATE within the PNG were the same).
- xnorswap 2y agoThat struck me as odd too. It may be just additional HTTP headers added to the response, but then it's hardly fair to use that as a point of comparison and treat smaller as "better".
- loeg 2y agoI think your guess is spot on. The QRcode images themselves are 594 and 577 bytes. The vast majority of the difference must be coming from other factors (HTTP headers). https://news.ycombinator.com/item?id=41973396 https://news.ycombinator.com/item?id=41973396
- pretzelhammer 2y agoAuthor here. The benchmarking tool I used for measuring response size was vegeta, which ignores HTTP headers in its measurements. I believe the difference in size is indeed in the QR code images themselves.
- jyap 2y ago
- bdahz 2y agoI'm curious what if we replace Rust with C/C++ in those tiers. Would the results be even better or worse than Rust?
- Imustaskforhelp 2y agoalso maybe checking out bun ffi / I have heard they recently added their own compiler
- znpy 2y agoIt should be pretty much the same. The article is mostly about exemplifying the various leve of optimisation you can get by moving “hot code paths” to native code (irrespective whether you write that code in rust/c++/c. Worth noting that if you’re optimising for memory usage, rust (or some other native code) might not help you very much until you throw away your whole codebase, which might not be always feasible.
- kelnos 2y agoIt should be about the same, though the main differences are likely to be caused by the speed of the QR code generator, and the PNG compressor. But assuming that the hypothetical C and C++ versions would be using generators and compressors of similar quality, it performance characteristics should be similar. The big plus(es) to using Rust over C/C++ are a) the C and C++ versions would not be memory-safe, and b) it looks like Rust's WASM tooling (if that's the approach you were to use) is excellent. (As someone who has written C code for more than 20 years, and used to write older-standard C++ code, I would never ever write an internet-facing server in either of those languages. But I would feel just as confident about the security properties of my Rust code as I would for my Java code.)
- lsofzz 2y ago<3
- Dowwie 2y agoBeware the risks of using NIFs with Elixir. They run in the same memory space as the BEAM and can crash not just the process but the entire BEAM. Granted, well-written, safe Rust could lower the chances of this happening, but you need to consider the risk.
- mijoharas 2y agoI believe that by using rustler[0] to build the bindings that shouldn't be possible. (at the very least that's stated in the readme.) > Safety : The code you write in a Rust NIF should never be able to crash the BEAM. I tried to find some documentation stating how it works but couldn't. I think they use a dirty scheduler, and catch panics at the boundaries or something? wasn't able to find a clear reference. [0] https://github.com/rusterlium/rustler https://github.com/rusterlium/rustler
- junon 2y agoI have no evidence of this but they may be liberally using catch_unwind: https://doc.rust-lang.org/std/panic/fn.catch_unwind.html https://doc.rust-lang.org/std/panic/fn.catch_unwind.html
- pjmlp 2y agoAnd so what we were doing with Apache, mod_<pick your lang> and C back in 2000, is new again. At least with Rust it is safer.
- ports543u 2y agoWhile I agree the enhancement is significant, the title of this post makes it seem more like an advertisement for Rust than an optimization article. If you rewrite js code into a native language, be it Rust or C, of course it's gonna be faster and use less resources.
- baq 2y ago'of course' is not really that obvious except for microbenchmarks like this one.
- ports543u 2y agoI think it is pretty obvious. Native languages are expected to be faster than interpreted or jitted, or automatic-memory-management languages in 99.9% of cases, where the programmer has far less control over the operations the processor is doing or the memory it is copying or using.
- baq 2y agoIt isn't obvious at all. A jit compiler has access to information that an aot compiler can only dream of. There aren't many languages which have both jit and aot compilers, though.
- ahoka 2y agoJava, C#?
- baq 2y agoyeah, that isn't 'many' and e.g. in java's case hotspot is a rather nice piece of engineering
- consteval 2y ago> A jit compiler has access to information that an aot compiler can only dream of If you know the machine and platform ahead of time, not really. For frontend JS this isn't the case. But for backend code it absolutely is the case. Sure, theoretically the JIT can sit in the background, see which functions are called the most and how they're call and then re-JIT pieces of code. In practice, I'm not sure how often this is done and if you even gain much performance. You MIGHT in a dynamically typed lang like JS because you can find out a bunch of info at runtime. In something like C# though? You already know a bunch at compile-time.
- djoldman 2y agoNot trying to be snarky, but for this example, if we can compile to wasm, why not have the client compute this locally? This would entail zero network hops, probably 100,000+ QRs per second. IF it is 100,000+ QRs per second, isn't most of the thing we're measuring here dominated by network calls?
- munificent 2y agoIt's a synthetic example to conjure up something CPU bound on the server.
- jeroenhd 2y agoWASM blobs for programs like these can easily turn into megabytes of difficult to compress binary blobs once transitive dependencies start getting pulled in. That can mean seconds of extra load time to generate an image that can be represented by maybe a kilobyte in size. Not a bad idea for an internal office network where every computer is hooked up with a gigabit or better, but not great for cloud hosted web applications.
- nemetroid 2y agoThe fastest code in the article has an average latency of 14 ms, benchmarking against localhost. On my computer, "ping localhost" has an average latency of 20 µs. I don't have a lot of experience writing network services, but those numbers sound CPU bound to me.
- demarq 2y agoI didn’t realize calling to the cli is that fast.
- kelnos 2y agoI doubt it's actually calling out to the CLI (aka the shell); presumably it's just fork()ing and exec()ing. On Linux, fork() is actually reasonably fast, and if you're exec()ing a binary that's fairly small and doesn't need to do a lot of shared library loading, relocations, or initialization, that part of the cost is also fairly low (for a Rust program, this will usually be the case, as they are mostly-statically-linked). Won't be as low as crossing a FFI boundary in the same process (or not having a FFI boundary and doing it all in the same process) of course, but it's not as bad as you might think.
- jinnko 2y agoI'm curious how many cores the server the tests ran on had, and what the performance would be of handling the requests in native node with worker threads[1]? I suspect there's an aspect of being tied to a single main thread that explains the difference at least between tier 0 and 1. 1: https://nodejs.org/api/worker_threads.html https://nodejs.org/api/worker_threads.html
- pretzelhammer 2y agoAs the article mentions, the test server had 12 cores. The Node.js server ran in "cluster mode" so that all 12 cores were utilized during benchmarking. You can see the implementation here (just ~20 lines of JS): https://github.com/pretzelhammer/using-rust-in-non-rust-servers/blob/main/node/server.js#L70-L92 https://github.com/pretzelhammer/using-rust-in-non-rust-serv...
- tialaramex 2y agoDoesn't "the 12 CPU cores on my test machine" answer your question ?
- Already__Taken 2y agoShelling out to a CLI is quite an interesting path because often that functionality could be useful handed out as a separate utility to power users or non-automation tasks. Rust makes cross-platform distribution easy.
- rwaksmunski 2y agoPretty sure Tier 4 should be faster than that. I wonder if the CPU was fully utilized on this benchmark. I did some performance work with Axum a while back and was bitten by Nagle algorithm. Setting TCP_NODELAY pushed the benchmark from 90,000 req/s to 700,000 req/s in a VM on my laptop.