5 ms·
SIMD on stable is very exciting news. SIMD unlocks the power of the GPU-esque parallelism that is already inside your CPU. While compilers do try very hard to t
by computerphage 8y ago
SIMD on stable is very exciting news. SIMD unlocks the power of the GPU-esque parallelism that is already inside your CPU. While compilers do try very hard to take advantage of this automatically, it’s not always predictable and can make performance fragile. This is what people are talking about when they refer to “low level control”. Don’t expect that with Rust 1.27 you’ll need to understand SIMD to get anything done. Do expect various libraries to just get faster on stable. For example, BurntSushi, the author of the regex crate and the ripgrep tool, has already enabled it if you’re on Rust 1.27 Stable. [1]
SIMD is another notch towards never needing to use nightly Rust. What this really looks like in practice is that less frequently will a user encounter a fast-mode and a slow-mode for a crate depending on whether they’re using nightly, with all the unstable features enabled, or not.
[1] https://github.com/rust-lang/regex/pull/490 https://github.com/rust-lang/regex/pull/490
- seanmcdirmid 8y agoHow useful is SIMD on CPU these days given that most of the touted original applications (back in the MMX SSE days) have been moved over to GPUs?
- vvanders 8y agoStill really useful. We use NEON on Android all the time since there's not a great API there and spinning up a GPU on a mobile device is not something you want to do unless you absolutely have to.
- flipgimble 8y agoI don't want to worry you, but on a mobile device with a screen the GPU is always spinning. That is where the pixels come from after all. I get your point that there are energy efficiency tradeoffs to consider. However for massively data parallel tasks like image operations, 3d graphics, machine learning, vision etc, I haven't seen SIMD implementation come close to GPU compute in terms of efficiency. Perhaps short lived text pressing and file compression where data access is more random could benefit from SIMD.
- vvanders 8y agoNope, I'm well aware. Built/worked on two android based 3D graphics stacks at two different companies and I've worked directly with just about every mobile GPU vendor aside from PowerVR(which is pretty similar to all the other tiled GPUs that are out there). Your GPU is actually running much less than you'd expect. Unless pixels are changing there's a 95% chance that everything on the system is sleeping and just the display controller is being kept spun up. Battery requirements mean that DVCS on any mobile chipset is going to be really aggressive. Either way GPU compute has a pretty large overhead both in terms of latency and scheduling. Earlier GPUs didn't schedule nice(hello waiting 16.6ms for your next compute request) and generally the type of places where you use SIMD(3D transforms, audio processing, etc) are so tightly coupled with other CPU operations that even the act of moving them to the SIMD registers is something you need to consider before diving into it. A lot of times waiting for some work queue to complete(or adding pipeline latency by waiting for next frame) just isn't feasible.
- jononor 8y agoThe screen of a mobile device is off just as much (or more) as it is on. More and more processing is happening in the background, when user is not actively doing something.
- stefan_ 8y agoPixels on screen come from the display controller, which on mobile is often a distinct piece of IP from the GPU, unlike desktop systems. Scenarios like watching a video full screen can often happen without GPU involvement at all.
- pcwalton 8y ago> We use NEON on Android all the time since there's not a great API there and spinning up a GPU on a mobile device is not something you want to do unless you absolutely have to. True regarding Android's APIs being awful, but taken literally, your statement implies that even Core Animation on-GPU compositing from 2007 is bad, which I'm sure you didn't mean. :)
- vvanders 8y agoYeah, this is all with in the context of compute, of course GPUs are going to be great at rasterizing(although they're rarely used for blitting/compositing different layers as there's discrete hardware that's even better for that).
- pjmlp 8y agoHave you tried to use Renderscript as well?
- vvanders 8y agoNo because RS abstracts away how the computation is queued, what latency is involved and has some pretty non-trival overhead in the framework its self. Like I mentioned above, earlier GPUs are pretty poor at fine grained scheduling + latency and NEON has pretty ubiquitous support on the target devices we were shipping.
- swsieber 8y agoWell, it seems pretty useful - e.g. it helps make ripgrep fast, it can speed up regex. There's a lot it can make faster. How much faster, I'm not sure.
- fulafel 8y agoGPUs are sadly not generally usable yet.
- blattimwind 8y agoVery useful. SIMD has a much lower barrier to use (doesn't need graphics drivers, GPGPU frameworks etc., almost universally available and fallbacks are easily implemented) and is much easier to target (same language, same toolchain, same memory model). Also notice that the execution model of GPUs and CPUs is quite different. You need a far larger "breadth" of execution to efficiently use a GPU, compared to a CPU.
- jedbrown 8y agoA GPU is much wider than a single core, but only slightly wider than a server CPU. For example, a 28-core Xeon has dual-issue FMA with 6-cycle latency and 16-wide packed SP registers, thus reaches peak floating point performance with 5376 independent operations in-flight at any instant. It's only about 4x higher for V100, which has higher TDP.
- makmanalp 8y agoStuff like compression algorithms, some codecs (not all use GPU), some high-performance parsers, some encryption stuff (e.g. openssl), some databases (often column stores, also redis I think), language VMs etc use SIMD. More generally SIMD is useful when you are repetitively performing the same instruction on a long stream of data but you don't want to send it over to the GPU because you don't want to incur the many order of magnitude slowdown of sending stuff back and forth to another chip or you can't rely on a GPU being there (e.g. embedded), or it's just overkill.
- majewsky 8y ago> [when] you can't rely on a GPU being there (e.g. embedded) Also, most servers have only very basic GPUs (if at all). Unless you're on a dedicated GPU server for machine learning etc., you're going to work with the CPU only.
- Entalpi 8y agoIn matemathics (and thus HPC/scientific computing/image processing/data science/scientific computing/etc) it is very common to multiply different matrices together which SIMD brings huge speedups for.
- _wmd 8y agoHeh. You can barely move your mouse across the screen without some layer in the architecture executing some SIMD instruction. It's everywhere, and it's going nowhere. Everything from counting the length of a string to rasterizing the text you're reading right now.
- pcwalton 8y agoI wish font rasterization used SIMD more, but it frequently doesn't. The final blitting step usually does, but important things like the Gaussian blur used for subpixel AA color defringing are still not accelerated on Mac, for instance :(
- jacobolus 8y agoI thought macs gave up on subpixel aa after the advent of retina displays.
- eridius 8y agoGood news, macOS isn't doing subpixel AA any more as of macOS 10.14! ;) But seriously, if SIMD would be useful there, why doesn't it use it?
- pcwalton 8y agoWell, I'm not Apple, so I can only speculate. I wouldn't be surprised if core low-level font rasterization isn't that well maintained, though. Often that stuff was written in the '90s and early aughts and so hasn't been touched.
- jcelerier 8y agoReal-time audio processing is bound to CPUs forever. It's the opposite of the data that works with GPU computations : very small chunks of data (generally between 64 and 512 floats or doubles) that you have to process sequentially. For those, SIMD is particularly useful - I got a ~7* increase in throughput recently when porting an effect from naive processing to AVX2 intrinsics - which means that artists can then put much more instances of the effect in parallel.
- mping 8y agoThere's a big cost of moving stuff in and out of the gpu. Unless it's a big/heavy workload, the CPU will be faster because of this.
- oldgeezr 8y agoA friend of mine is playing with neural networks and training them to play reversi. He's working with a lot of matrices so he tried AVX extensions and CUDA. AVX runs circles around CUDA, probably because of the setup time of moving things back and forth to the GPU. Also, CUDA can be a big pain in the ass to get working.
- stochastic_monk 8y agoAlso, if your code has a lot of branching (most of my work wouldn’t benefit from offloading to GPU), or if the data being processed in parallel at a time is too small to make up for memory transfer, it can be the right approach and provide a huge performance boost.
- mrbonner 8y agoI recently ported some code to use AVX intrinsics. The performance gain is somewhere from 30x to 3000x depending on the size of the array. The larger the array the more performance gain I get. It’s linear algebra but not related to matrix multiplication that I could mention.
- eternalban 8y agoCorrect my possible misunderstanding but isn't this fabled "power" at best x4 performance gain and not really in the same league as GPU-esque?
- antoinealb 8y agoDepending on your operations, e.g. With small buffers that fit in a cache line, transfers to GPU memory could cause GPU esque performance of less than 4x, even less than 1x whereas SIMD still provides benefits.