15 ms·
Towards fearless SIMD, 7 years later
- the__alchemist 2y agoVery interesting! I posted a vector and quaternion lib here a few weeks ago, and got great feedback on the state of SIMD on these things. I since have went on a deep dive, and implemented wrapper types in a similar way to this library. Used macros to keep repetition down. Split into three main sections: - Floating point primitives. Like this lib. Basically, copied `core::simd`'s API. Will delete this part once core::simd is stable. `f32x8`, `f64x4` types etc, with standard operator overloads, and utility methods like `splat`, `to_array` etc. - Vec and Quaternion analogs. Same idea, similar API. Vec3x8, Quaternionx8 etc. - Code to convert slices of floating point values, or non-SIMD vectors and quaternions to SIMD ones, including (partial) handling of accessing valid lanes in the last chunk. I've incorporated these `x8` types into a WIP molecular dynamics application; relatively painless after setting up the infra. Would love to try `Vec3x16` etc, but 512-bit types aren't stable yet. But from Github activity on Rust, it sounds like this is right around the corner! Of note, as others pointed out in the thread here I mentioned, the other vector etc libs are using the AoS approach, where a single f32x4 value etc is used to represent a Vec3 etc. While with this SoA approach, a `Vec3x8` is for performing operations on 8 Vec3s at once. The article had interesting and surprising points on AVX-512 (Needed for f32x16, Vec3x16 etc). Not sure of the implications of exposing this in a public library is, i.e. might be a trap if the user isn't on one of the AMD Zen CPUs mentioned. From a few examples, I seem to get 2-4x speedup from using the x8 intrinsics, over scalar (non-SIMD) operations.
- camel-cdr 2y agoWhy do the x4/x8 types seem to be the default in rust? A portable SIMD feature should encurage portable SIMD and not a specific vector register size.
- the__alchemist 2y agoThis sounds like a great idea. I went with this approach because I'm new to SIMD, so I aped the most promising API (core::simd), extending it naturally. I need to think through the consequences. It might involve feature gates, and/or an enum. So, for example, instead of: pub struct f32x8(__m256); It might be this internally, with some method to auto-choose variant based on system capability?: pub enum f32_simd { X8(__m256), X16(__m512), } etc. Thoughts?
- camel-cdr 2y agoI'm not proficient in Rust, but API wise I'd conceptually define types like f32xn or f32s, which have the number of elements that fit into a vector register for your target architecture, so 4 for NEON/SSE, 8 for AVX and 16 for AVX512. I can recommens lookig at the highway library: https://github.com/google/highway https://github.com/google/highway
- janwas 2y agoAmen! I really do not understand this. It has been 7 years since SVE was introduced. Writing an application in terms of a specific lane count loses performance portability - either it's too many, or too few, for the particular CPU. And it also enables/encourages antipatterns like putting RGB in Vec4.
- dzaima 2y agoSeems rustc nightly does successfully vectorize the first sigmoid example: https://rust.godbolt.org/z/e1WYexqWY https://rust.godbolt.org/z/e1WYexqWY Also there's progress on making safe intrinsics safe: https://github.com/rust-lang/stdarch/pull/1714 https://github.com/rust-lang/stdarch/pull/1714
- ashvardanian 2y agoI’ve said it before and I’ll say it again: Rust feels like a Python developer’s idea of a high-performance computing language. It’s a great language for many kinds of applications — just not when you need to squeeze out every bit of performance from advanced hardware. Even before getting into SIMD, try using Rust for concurrent, succinct, or external-memory data structures. It quickly becomes clear where the friction is. Cargo is fantastic — clean, ergonomic, and a joy compared to many toolchains. But it’s much easier to keep things simple when you don’t have to support dozens of AVX-512 variants, AMX, SME, different CUDA generations, ROCm, or any of the other modern hardware capabilities. Standardising SIMD in the standard library — in Rust or C++ — has always been a questionable idea. Most of these APIs cater to operations that compilers already auto-vectorize reasonably well, and they barely touch the recent capabilities of SIMD. Just consider how hard it is to build any meaningful abstraction over the predicate/register models across AVX-512, SVE, and RVV. RVV aside, this should illustrate the point: https://www.modular.com/blog/understanding-simd-infinite-complexity-of-trivial-problems https://www.modular.com/blog/understanding-simd-infinite-com...
- lifthrasiir 2y ago> Just consider how hard it is to build any meaningful abstraction over the predicate/register models across AVX-512, SVE, and RVV. Note that Highway mentioned in the post does take care of this, which is no easy feat but also a proof that it is doable.
- the__alchemist 2y agoI guess it comes down to application. If you don't attempt to find the most general solution, you can dodge those pitfalls. Case in point, abstracting over AVX-512, SVE, and RVV may be tough, but picking one is fine (On nightly only for now), can with the right abstractions can be almost as straightforward as using normal scalar values. I don't have a solution on the CUDA variants either; have been hard-coding that as well... (Cudarc lib with CUDA-version feature gates and GPU-series-specific code). Haven't hit a brick wall yet, but might... or might not.
- dzaima 2y agoI don't think Rust is particularly problematic here. As long as you don't want to do funky things like use immutable argument memory as temporary scratch space (with you restoring the values afterwards of course), all it means is some `unsafe`ing at worst, compared to C/C++. And there are some safe abstractions you can make over loads/stores (everything else being safe, even if not yet marked as such). Do agree that a standard SIMD type is rather pointless, if not immediately, then in like 5 years. (and, seemingly, both Rust and C++ are like over 10 years behind on SIMD, so they're already out-of-date) Maybe somewhat useful if you just want the simple ~8x speedup, and not squeeze out the last 1.4x or whatever, but autovectorization should be capable of covering a significant amount of such.
- curtisszmania 2y ago[dead]
- isusmelj 2y agoI’ve been playing around with SIMD since uni lectures about 10 years ago. Back then I started with OpenMP, then moved to x86 intrinsics with AVX. Lately I’ve been exploring portable SIMD for a side project where I’m (re)writing a Numpy-like library in Rust, mostly sticking to the standard library. Portable SIMD has been super helpful so far. I’m on an M-series MacBook now but still want to target x86 as well, and without portable SIMD that would’ve been a headache. If anyone’s curious, the project is here: https://github.com/IgorSusmelj/rustynum https://github.com/IgorSusmelj/rustynum. It's just a learning exercise for learning Rust, but I’m having a lot of fun with it.
- IshKebab 2y agoA problem for RISC-V is going to be that there's currently no way for user code to detect the presence of RVV. I have no idea how you can do multiversioning with that limitation.
- hmry 2y agoThe solution is to ask the OS to detect it for you. Linux offers a syscall for this (riscv_hwprobe). Has the drawback that it requires OS support, of course. But RVV requires OS support anyway (e.g. managing mstatus, saving vector registers on context switch), so that seems fine to me.
- dzaima 2y agoThere is some work on an OS-agnostic feature detection C API: https://github.com/riscv-non-isa/riscv-c-api-doc/blob/main/src/c-api.adoc#extension-bitmask https://github.com/riscv-non-isa/riscv-c-api-doc/blob/main/s.... Still quite new though, and potentially might change (as it did a month ago).
- janwas 2y agoOr also getauxval? Highway has code to check for this, including that vectors are at least 128 bits: https://github.com/google/highway/blob/master/hwy/targets.cc#L664 https://github.com/google/highway/blob/master/hwy/targets.cc...
- thomashabets2 2y agoOh? Isn't that what this does? std::arch::is_riscv_feature_detected!("v") Hmm… now that I actually experiment with it, I can't get it to return `true` on hardware that does support it, unless I also compile with -Ctarget-feature=+v. And if I do, then the binary crashes with SIGILL before getting to that point on hardware without rvv. So if it's always equal to cfg!(target_feature="v"), then what does that even mean? I have created https://github.com/rust-lang/rust/issues/139139 https://github.com/rust-lang/rust/issues/139139
- fulafel 2y agoWhat happens there when you try to execute a missing RVV instruction ? On other archs you get a SIGILL which you can handle.
- DeathArrow 2y agoUsing SIMD in C#: https://xoofx.github.io/blog/2023/07/09/10x-performance-with-simd-in-csharp-dotnet/ https://xoofx.github.io/blog/2023/07/09/10x-performance-with... https://learn.microsoft.com/en-us/dotnet/standard/simd https://learn.microsoft.com/en-us/dotnet/standard/simd