7 ms·
This looks kinda gross to me. Do the rust developers not want to emulate what ipsc and cuda do? Writing intrinsics by hand is not what I expect from a 2019 lang
by krapht 7y ago
This looks kinda gross to me. Do the rust developers not want to emulate what ipsc and cuda do? Writing intrinsics by hand is not what I expect from a 2019 language.
- psv1 7y agoTo be honest, everything in Rust looks a bit ugly to me. I really tried to like the language but the syntax, everything being overly annotated, and the number of features that you need to understand to do simple tasks - all of these make it not really worth it to pick Rust for new projects. There are other problems like the ecosystem of crates and the lack of learning resources but at least they aren't intrinsic to the language itself.
- ekidd 7y ago> To be honest, everything in Rust looks a bit ugly to me. I write a lot of Rust code at work, and I admit that it can sometimes be pretty noisy. There are several major contributors to this: 1. Rust offers fine-grained control over pass-by-value, pass-by-reference, and pass-by-mutable reference. This is great for performance. But it also adds a lot of "&" and "&mut" and "x.to_owned()" clutter everywhere. 2. Rust provides support for generics (aka parameterized types). Once again, this is great for performance, and it also allows better compile-time error detection. But again, you wind up adding a lot of "<T>" and "where T:" clutter everywhere. 3. Usually, Rust can automatically infer lifetimes. But every once in a while, you want to do something messy, and you end up needing to write out the lifetimes manually. This is when you end up seeing weird things like "'a". But in my experience, this is pretty rare unless I'm doing something hairy. And if I'm doing something hairy, I'm just as happy to have more explicit documentation in the source code. Really, the underlying problem here is that (a) Rust fills the same high-control, high-performance niche as C++, but (b) Rust prefers explicit control where C++ sometimes offers magic, invisible conversions. (Yes, I declare all my C++ constructors "explicit" and avoid conversion operators.) Syntax is a hard problem, and I've struggled to get syntax right for even tiny languages. But syntax for languages with low-level control is an even harder problem. At some point, you just need to make a decision and get used to it. In practice, I really enjoy writing Rust. It's definitely not as simple as Ruby, Python or Go. But it fills a very different ecological niche, with finer-grained control over memory representations, and support for generics.
- gameswithgo 7y agothis is exactly how simd intrinsics look in c, its not rust thing its an intrinsics thing.
- vardump 7y agoIntrinsics is still what modern high-performance "C++" code uses. Auto-vectorizers are pretty fragile and require too much babysitting to be worth it.
- The_rationalist 7y agoSemi auto vectirizers are the best compromise Cf openMP and also allow to multithread or offload your code on the gpu.
- vardump 7y agoSometimes offloading to GPU is not possible or desirable for a reason or another. Like total latency, no point to offload something you can finish processing on CPU faster than transferring to GPU and back. Some systems just don't have GPUs, and there's nothing you can do about it. Sometimes CPUs are simply much faster due to a branchy serial algorithm. However, you might still be able use SIMD to get some speedup. Sometimes I end up going single threaded SIMD, if the whole system is memory bandwidth limited anyways. Work stealing queues can also be great. Thread per CPU core pulling work from a common pool. You might be able to do some rough data locality based scheduling to reuse cache hierarchy contents. Overall, I feel the biggest challenges often come from cache and memory bandwidth management. CPUs are fast, but SDRAM is not. You don't want different threads fighting for CPU socket local resources and even less for global ones. I usually do rough estimates of required bandwidth and computation, write some prototypes and do a lot of profiling, including taking a good hard look at the CPU counters. Not trying to say anything particular, except that solution space has some options. That there are no silver bullets. The solutions you suggested can also be great.
- dragontamer 7y agoIPSC gives you SIMD code on the CPU, but programs as if the CPU-SIMD units were a GPU. Its an excellent project. --------- > Overall, I feel the biggest challenges often come from cache and memory bandwidth management. CPUs are fast, but SDRAM is not. You don't want different threads fighting for CPU socket local resources and even less for global ones. I usually do rough estimates of required bandwidth and computation, write some prototypes and do a lot of profiling, including taking a good hard look at the CPU counters. I think memory-layout is the #1 issue these days. CPUs / GPUs have so much compute available that its almost impossible to actually achieve high utilization. In most cases, you're sitting around just waiting for memory... CPU memory movement is still subpar compared to GPUs. AVX512 finally implements "scatter" operations, but GPUs have had highly-optimized "gather-scatter" to __local or __shared__ memory for years (ex: GPUs have 32 banks and 32-load/store units per GPU-compute unit or NVidia SM: that's either 1/2 or 1 load/store unit per GPU shader. AVX512 Skylake however has 3-load/store units across 16 SIMD-threads...) Intel really needs to write more instructions like "pshufb" to handle more ways for register-to-register movement. It seems like a lot of data-movement in the AVX world is still best handled by AVX -> L1 cache -> back into AVX register (which is limited by the very few load/store units in modern CPUs). Yeah, you can cheat a lot of cases through pshufb, but that instruction doesn't always work. There's something to be said about the brute-force option of 32x load/store units on a GPU-unit and sticking 32-load/store units for all the threads to leverage.
- maeln 7y agoThis is an example of doing manual SIMD. In some cases, it's still needed. Manual SIMD tend to be faster than what even the most modern compiler can provide. Because it is low level, it won't be fancy, but you can find several library that wrap those low-level ops in more fancy APIs. Comparing it with CUDA is hardly a good comparison. Even if the GPU is basically a bunch of SIMD unit, GPGPU programming is still very different than adding SIMD capability to an x86 program.
- dragontamer 7y agoYou really should check out ipsc, which is a CPU-SIMD compiler similar to CUDA. With that being said, there's good reason to use raw intrinsics in modern code. But ipsc / CUDA model is superior for most uses in my experience. Its just easier to think about. The main issue with IPSC is that you're innately SOA, and the data-layout is just different compared to how people normally organize their data. Data-layout issues (AOS vs SOA) are probably one of the most tedious issues to deal with when using SIMD. For the "interface", where you're converting AOS to SOA, manual intrinsics can help.
- sgift 7y agoThere are libraries abstracting the SIMD calls if that's what you want, e.g. faster (https://github.com/AdamNiederer/faster https://github.com/AdamNiederer/faster) or simdeez (https://github.com/jackmott/simdeez https://github.com/jackmott/simdeez), but these have to use basic operations at the end of the day too and this post shows how to do that. I'm not sure renaming the primitive operations provided by intel/amd to something "nicer" would help much here. Using plain SIMD will always be ugly and at least you can Google the names and get back the Intel documentation without first translating from a different name.
- dragontamer 7y agoI disagree with the parent poster's phrasing, but they have a point: CUDA (for GPUs) and IPSC (for CPUs) makes SIMD code far easier to write. I think OpenMP / Intel Autovectorizers / etc. etc. are all taking the wrong approach. The "graphics guys" have figured out a better model for thinking in SIMD. With that being said, normal code has major issues before it can be converted into "Graphics-SIMD" form. Most importantly: data-layout is straight up "wrong", with most data-layouts in AOS (array of structs) instead of SOA form. Writing code that interfaces between AOS and SOA is tedious, and I'm unconvinced that any general solution can be done automatically. (Remember: the key is to convert between the forms "efficiently", because the only reason we're putting up with SIMD at all is due to performance reasons).
- fluffything 7y ago> Do the rust developers not want to emulate what ipsc and cuda do? Of course, which is why the language already allows you to do that if you want to, and it often performs better than ISPC, while being memory and thread safe. However, because Rust is a low-level language, it also allows you to write low-level code that uses assembly-like intrinsics for specific instructions manually, which is what this blog post shows. I personally think that if your goal is to teach a new programming paradigm, like data-parallel programing (SIMD/SIMT/..), using assembly is a pretty inefficient way to do that. If you already know a data-parallel programming language like ISPC, then there is a lot of value on learning to which assembly instructions your code should lower to on each hardware and getting an intuition for that.
- BubRoss 7y agoPerforming better than ISPC is a pretty bold claim, you are definitely going to need to provide a source for that. Rust developers have had people telling them about ISPC for years and have waived off any need to look at it and understand why it works so well.
- fluffything 7y agoSee https://github.com/rust-lang-nursery/packed_simd#performance https://github.com/rust-lang-nursery/packed_simd#performance Apparently the results have a spread, depending on the CPU used to perform the benchmark - they appear to be testing from haswell to skylake so that covers a wide range of x86 hardware. Matching ISPC perf in the low end, and being ~1.5x faster on the high end, while being able to target all hardware that Rust can target (arm, wasm, ppc, riscv...) sound better than ISPC to me, which works for x86 very well, but not so good for ARM, and not really at all for anything else.
- The_rationalist 7y agoRust is lagging behind in the performance world, especially because it lack openMP/ACC support.
- gameswithgo 7y agoThe rayon library provides a lot of the same features as those.
- epage 7y agoSome further background: vendor instrinsics are well documented and unchanging. They are a lot easier for a language to standardize on. As others have pointed out, there are higher level libraries being built on top of them. This gives the Rust community the freedom to experiment on how the instrinsics should be implemented with less concerns over compatibility. This also is a nice way of handling a limited subset of other assembly instructions for systems programming while they figure out how to have inline assembly without coupling the language to its implementation.