4 ms·
Someone posted a great question here that they later deleted. My full response also wasn't submitted (and lost), so I'm trying to recreate part of it here. It w
by firebones 10y ago
Someone posted a great question here that they later deleted. My full response also wasn't submitted (and lost), so I'm trying to recreate part of it here. It was actually a good question about whether Rust's efficiency has been tested.
I've done some testing; Java is faster in terms of the sip test for a lot of stuff, especially for cases where you micro-benchmark single processes without regard to overall CPU or memory usage. However, when you scale up (to the point where JVM tuning comes into play), the raw performance in terms of response time is comparable, while the overall system utilization (e.g., "time" measurements of user/sys) is much better for Rust--Java uses a lot more system time due to JVM housekeeping threads. So more efficient is not a win for Java, even though wall-clock benchmark time might look better for a single process (as long as you don't look at overall system utilization.)
Java's performance comes at the price of more memory usage and higher variance in response time. You can attempt to constrain one or the other, but you can't constrain both at the same time. Microbenchmarks typically put a high enough ceiling on memory that they're relevant for very constrained loads. To be robust beyond the microbenchmark, incremental GC is required, which levels the playing field between Java GC and Rust's jemalloc. Still, even when you attain parity, the Java solution is using far more memory due to the JVM's OOP overhead.
Another example of a hidden "cost" of Rust comes from things like Unicode support. Recently I tested a regex-heavy algorithm that used burntsushi's regex library for Rust. In initial testing, Java blew it away. What I later realized was that the default Java implementation I used did not support Unicode characters. When I enabled that support, and enabled incremental GC (to support the scale of testing I was performing), the performance was similar.
This is another example of Rust being mischaracterized up front. Older languages took short cuts, and fare well for the low end of testing (the microbenchmark). Rust tends to look forward and to the bigger picture.
Anyway, sorry you deleted your legitimate question. It was a good one, and the kind that will make Rust better.
- Noseshine 10y ago(OT) > My full response also wasn't submitted (and lost), so I'm trying to recreate part of it here. Just an aside, that's exactly why I installed browser extension "Lazarus: Form Recovery" years ago. Let's me easily recover/reuse anything I typed into a form previously. Hasn't been updated since 2014 (Chrome web store) but works fine.
- burntsushi 10y ago> Another example of a hidden "cost" of Rust comes from things like Unicode support. Recently I tested a regex-heavy algorithm that used burntsushi's regex library for Rust. In initial testing, Java blew it away. What I later realized was that the default Java implementation I used did not support Unicode characters. When I enabled that support, and enabled incremental GC (to support the scale of testing I was performing), the performance was similar. Could you explain a bit more about this? I find it surprising. If you can't share the code, perhaps you could share the regexes? Which Java regex engine did you use? (There really should only be one case where Unicode support causes performance problems, and that's when you use word boundaries.)
- firebones 10y agoIt was a find_iter() across \w+. There was other surrounding code that might have affect the output (it emitted (String, position pairs). I will try to isolate a test case and reach out... BTW, your fst is great stuff.
- burntsushi 10y agoThanks! If you come up with an example I'd love to see it. Generally, even though `\w` in Rust's regex library supports Unicode, it shouldn't result in a slow-down compared with the non-Unicode `\w`, assuming you're using find_iter. (Of course, Unicode support isn't free, but the primary cost here is memory and compile time, not matching performance.) If you were indeed emitting `String` (a new allocation for every match) instead of `&str`, then that could certainly be a possible explanation for the slow down.