4 ms·
On the last paragraph, for runtime selection of SIMD implementations, would that use something like STT_GNU_IFUNC? Perhaps exactly that on Linux, since it's co
by CUViper 11y ago
On the last paragraph, for runtime selection of SIMD implementations, would that use something like STT_GNU_IFUNC? Perhaps exactly that on Linux, since it's compiling to ELF, after all.
I would have liked to see benchmarks compared to non-simd rust too. Thankfully this is in the source code. Here's what I get from cargo bench, fwiw:
Running target/release/mandelbrot-aefed80dbc3f2841
running 2 tests
test mandel_naive ... bench: 802,072 ns/iter (+/- 25,106)
test mandel_simd4 ... bench: 235,853 ns/iter (+/- 6,374)
test result: ok. 0 passed; 0 failed; 0 ignored; 2 measured
Running target/release/matrix-58de8ccd4bd58dcd
running 6 tests
test inverse_naive ... bench: 4,967 ns/iter (+/- 251)
test inverse_simd4 ... bench: 1,984 ns/iter (+/- 94)
test multiply_naive ... bench: 2,226 ns/iter (+/- 26)
test multiply_simd4 ... bench: 897 ns/iter (+/- 32)
test transpose_naive ... bench: 627 ns/iter (+/- 16)
test transpose_simd4 ... bench: 361 ns/iter (+/- 7)
test result: ok. 0 passed; 0 failed; 0 ignored; 6 measured
- dbaupp 11y agoYeah, the runtime selection could indeed leverage ifunc when available (the actual dispatch isn't so interesting, since a simple branch on CPUID is perfectly workable in many cases: usually the dispatch is done when calling an expensive, long-running function, so a little bit of O(1) setup is in the noise). The really hard bit is getting dependencies compiled in multiple modes using different features, which will be especially important if/when Rust gets an ecosystem of libraries with SIMD utilities/functions. I hope we do get a wide variety of such things (rather than everyone just using the intrinsics directly, as in C/C++). BTW, the benchmarks are comparing to non-SIMD rust: the graphs are of how many times faster the SIMD Rust code is than scalar Rust code (i.e. if they were plotted, the scalar bars would all be at 1.0).
- CUViper 11y agoAha, the y-axis is "times faster than scalar", I totally missed that.