3 ms·
A few of the issues: 1) Not much code is written to take advantage of them. This seems like inherently less of a problem with SVE, assuming it becomes widely a
by celrod 6y ago
A few of the issues:
1) Not much code is written to take advantage of them. This seems like inherently less of a problem with SVE, assuming it becomes widely adopted, as this should naturally support compiling source to take advantage of whatever the vector width of the host CPU happens to be. In the same vein, SVE easily supports market segmentation without the instruction set extension fragmentation that's been causing problems for Intel.
The Arm SIMD Implementations lists the following SVE cpus:
- Neoverse N2 (2x 128 bit units, comparable to Atom w/ SSE only)
- Neoverse V1 (2x 256 bit units, comparable to AMD's offerings and Intel's desktop offering, w/ AVX2)
- A64FX (2x 512 bit, comparable to Intel server and HEDT with AVX512 [not all of which have 2x fma units]).
For the x86_64 CPUs, you'd need three binaries or function multiversioning to support them all and make the most of the hardware. With ARM, you could do it with a single ordinarily compiled binary.
2) Downclocking, already mostly solved on Ice Lake: https://news.ycombinator.com/item?id=24215022 https://news.ycombinator.com/item?id=24215022
Using more of the chip will obviously use more power, so of course you'll have to hit thermal or power limits sooner. But wider units are more efficient, so you'll get more work done before hitting such limits; the problematic downclocking was where a handful of such instructions would trigger it, out of proportion of energy requirements.
3) Not all workloads can benefit from wide SIMD units. Some people (famously Linus Torvalds) would be much better off and happier with using the space for more cores.
As for benefits, if you like HPC-like workloads (numerics that are hard to offload to the GPU), wide SIMD units on the CPU can be an excellent way to accelerate it. This benefits that audience most.
- exged 6y agoSVE doesn't really solve #2 and #3 by itself. You can still design a core with huge SIMD units that require a ton of power and area, resulting in downclocking and relatively poor performance on non-SIMD workloads. #1 is a big deal though. It's a huge burden to rewrite SIMD code over and over for every new instruction set, so people (and compiler implementers) just don't use it too often. Then you're paying power and die area for completely useless SIMD units.
- vlovich123 6y agoI don't understand how #1 is possible. How do you statically compile the code once & have it automatically choose the right vector width that's available on the CPU at run time? Does the compiler just emit an instruction saying "ideal width at this point is 256bits" & the CPU will automatically optimize that however needed? I've been unable to find any meaningful description of how this actually works.
- jedbrown 6y agoAn example showing how one writes vector-length agnostic code using C intrinsics, and the assembly that is produced: https://developer.arm.com/documentation/100891/0612/coding-considerations/using-sve-intrinsics-directly-in-your-c-code https://developer.arm.com/documentation/100891/0612/coding-c...
- vlovich123 6y agoCould just be me, but that seems mind-bogglingly complex vs regular vector instructions & could easily make some traditional programs even harder to vectorize.