6 ms·
Some of the rust versions calls C libraries for its heavy lifting (gmp, pcre) so I wouldn't take this too seriously.
by mh7 6y ago
Some of the rust versions calls C libraries for its heavy lifting (gmp, pcre) so I wouldn't take this too seriously.
- sitkack 6y agoAs soon as the libraries are RiiR, then the Rust compiler can optimize across those library calls.
- burntsushi 6y agoThat's not good enough. Rust already has a pure-Rust regex library. (I'm its author.) It is the only non-PCRE regex engine to appear in the first 20 results of the regex-redux benchmark. (The 21st I believe is currently RE2.) When using the regex crate, you would not materially benefit from optimizations across library calls, nor is it the difference maker here. Highly optimized regex engines depend more on internal inlining. Take a look at the object code for a program compiled with PCRE2 or the regex crate. You'll find huge functions internal to the regex library where inlining has been forced to reduce overhead. Those things are never going to be inlined across library boundaries.
- sitkack 6y ago> Those things are never going to be inlined across library boundaries. What prevents this? I trust you on on this, but where is the remaining work? Language semantics, compiler, third choice?
- burntsushi 6y agoIt's prevented by good sense. The functions are likely multiple KB in size. Inlining them would seriously bloat the binary and would be unlikely to help due to how much work most regex engines do on each search. The remaining work _on this particular benchmark_ is the regex algorithm itself. I'm on mobile so I can't do a deep dive, but I haven't yet figured out how to easily improve on this particular case. It has to do with the fact that the benchmark has a high match count and the finite automata approach in the regex crate has a bit higher overhead than the typical backtracking solution used in PCRE2 (which is also JIT'd in this case). It's not the language, compiler, inlining or any other such thing. It's algorithms. But this is one single benchmark. Before regex-redux there was regex-dna, and Rust's regex crate was #1 there. Why? Same reason. Algorithms. You can't judge regex performance by a single benchmark. Two won't do it. Not even ten. It's one of the many problems with the Benchmark Game. This would be fine if everyone was circumspect and careful with their interpretation of the data provided, but they aren't. And the Benchmark Game doesn't really do much to alleviate this other than some perfunctory blurbs in some corners of the web site. With that said, running a benchmark is hard work. It's easy to criticize.
- igouy 6y ago> … doesn't really do much to alleviate this … It's easy to criticize. Indeed. Especially easy if it's just fault-finding without any suggestion as to what might be done to "alleviate this" :-)
- burntsushi 6y agoYou can do better than that, but I suppose I don't expect more than an insubstantial pithy quip from you. Add analysis to each benchmark. Require submissions to come with analysis. Or more minimally, make the existing disclaimers on the web site more discoverable through one of any number of means, up to taste.
- igouy 6y ago> Add analysis to each benchmark. Perhaps you'd like to take on that task? > Require submissions to come with analysis. Which raises the barrier for program contributors and presumably would require me to judge whether their analysis was acceptable? Me? Really? ;-) > disclaimers on the web site more discoverable The problem is that no one wants to "discover" disclaimers. We want to see something that supports whatever it is we already believe — and we're very good at ignoring anything else.
- burntsushi 6y agoYou asked me what could be done. I answered. I'm sure you have your reasons why you don't do more. But you asked. > The problem is that no one wants to "discover" disclaimers. Again repeating something I already said. It is indeed a problem. There are ways to address it. I personally think you do the absolute minimum.
- igouy 6y ago> It is indeed a problem. There are ways to address it. Please make a specific suggestion.
- pjmlp 6y agoLooking at the progress of Rust/WinRT, and C++ renaissance thanks to GPGPU computing and mobile OSes stack, that is still a couple of decades away. And then there are the whole LLVM and GCC based eco-systems.
- ncmncm 6y agoWhen a Rust program is faster than the matching C program, it is utterly nonsensical to attribute its speed to C. It is gratifying to see C++ identified, here, as the hands-down fastest implementation language, but odd to see Rust performance still compared, in the headline, to C, as if that were the goal. The headline should say that Rust speed is approaching C++ speed. In principle, Rust speed should someday exceed C++'s, in some cases, because it leaves behind some boat anchors C++ must retain for backward compatibility. In particular, if compiler optimizers could act on what they have been told about language semantics, they should be able to do optimizations they could not do on C++ code. As it is, Rust relies on optimizers coded for C and C++. Rust may never reliably beat C++, because Rust has itself committed to details that interfere with optimization, principally in its standard library: The C++ Standard Library offers more knobs for tuning, while the Rust libraries are simpler to use. Some of the reasons that C++ is so reliably faster than C are subtle and, to some, surprising. As noted elsewhere in this thread, the compiler knows more about what the language is doing, and can act on that knowledge, but C and C++ compilers share their optimizer, so that makes less difference than one might guess. Mainly, the C++ code does many things you might do in C, but in exactly one way that the optimizer can recognize easily. The biggest reason why C++ is so much faster than C is that C++ can capture optimizations in libraries and reliably deliver those optimizations to library users. The result is that people who write C++ libraries pay a great deal of attention to performance because it pays, and that attention directly benefits all users of the library, including other libraries. Because you can capture more semantics in C++ libraries, C++ programs naturally use better algorithms than C programs can. C programs much more frequently use pointer-chasing data structures because those are easier to put into a C library, or quicker to open-code, where the corresponding C++ program will use a library that has been hand-tuned for performance without giving up anything else. Rust gets to claim many of the same benefits, because it is also more expressive than C, and Rust libraries are often as carefully tuned. Rust is not as expressive as C++, yet, and has many, many fewer people working to produce optimal libraries, but it is doing well with what it has.
- pjmlp 6y agoI agree, however from my experiences with lifetime checkers in VC++, it will still take a couple of more years to make it work properly and hardware memory tagging as being pushed by Oracle, Microsoft, Apple, Google and ARM is still a couple of years away to be deployed everywhere.