4 ms·
Looks like the author is going to enjoy RISC-V SIMD (once it is finalized).
by devit 8y ago
Looks like the author is going to enjoy RISC-V SIMD (once it is finalized).
- glangdale 8y agoI'm not sure... I like the look of the bit manipulation stuff, but I'm not really a fan of the variable-length vector approach. I think these systems are built to make the world safe for matrix multiply - and simple math workloads - but the short-vector approach (e.g. the typical x86 SIMD style of doing things) is sometimes exactly what you want. I keep meaning to write more about this. I have found 3 categories of SIMD use in my own work: 1) Doing one thing a gazillion times (e.g. conventional SIMD). This works well on vector machines, of course. 2) Using SIMD registers to do "more stuff". So in a couple string matchers and regex matchers I've designed, you are really use using a SIMD register because it's bigger than a GPR register (duh). But this might be because you want to simulate a 512-bit NFA rather than a 64-bit NFA. 3) Using SIMD operations to do weird, irregular stuff where what you're really getting is a substitute for branchy code. I blogged about an example of this called "SMH" https://branchfree.org/2018/05/30/smh-the-swiss-army-chainsaw-of-shuffle-based-matching-sequences/ https://branchfree.org/2018/05/30/smh-the-swiss-army-chainsa...
- mastax 8y agoI'll grant that a lot of code can't be readily adapted to large vectors, but this doesn't seem like much of a problem for (proposed) RISC-V. If you're creating an algorithm that only works on a specific width, you can just `setvl` and assert that you have enough space available. You're likely targeting a specific CPU or class of CPU where you know there will be support for 256-bit vectors or whatever. If someone tries to run it on a cheap embedded CPU, they'll be disappointed that your code won't run on their 64-bit vectors, but this is no different from trying to run AVX code on an Intel Atom. I suppose it's hard to know until we can actually write code for it, but I haven't imagined a scenario where the RV model is significantly worse than packed SIMD, other than a tiny bit of vector configuration bookkeeping. I think that's a worthwhile tradeoff for getting simple, portable, fast code for elementwise operations and implementation flexibility. I'd love to be convinced otherwise, though.
- jcranmer 8y agoThere's really two kinds of vector code in my experience. The first kind is the standard vector code that most people probably think of, the loops that boil down to: #pragma vectorize_me_to_death for (int i = 0; i < N; i++) { // ... } But there is also a second kind of vector opportunity: small vector opportunities. SLP vectorization is the ur-example here: you scan a block of code for operations that happen to be doing the same operation on different values and make vector code out of it. For this kind of vector code, there is a lot more focus on horizontal and shuffling code than the wide kind of vector code. I haven't looked at the ARM SVE or RISC-V vector ISAs in detail, but I imagine that they don't support the latter kind of vectorization very well.
- brandmeyer 8y agoI'm not sure I understand how callee-saved registers are going to work under RVV, given the way that dynamic reconfiguration works.
- devit 8y agoMy guess is that either no registers will be callee-saved or one group of 8 registers will be callee-saved (in the current draft, you can group 1, 2, 4 or 8 sequential registers, but only starting on a multiple of the group size, so register groups are always going to be contained within one of the four 8-register maximum-size groups); this will require dynamic stack allocation since the vector length is not fixed. The saving sequence would be something like this (after setting up a frame pointer if needed): "vsetvl t0, x0, e8, m8; sub sp, sp, t0; vse.v v16, (sp)".
- glangdale 8y agoThe pain of variant SIMD sizes is real, I admit. In the Intel world this means 3 sizes at the moment: 128-bit for Atom and really old stuff (if you care), 256-bit for most mainstream processors and 512-bit for cutting edge - then pick your baselines _within_ those. Not fun. On the other hand, we can compare the current ARM situation, where everyone is stuck at 128 for now. A "tiny bit of vector configuration bookkeeping" has my hackles going up. I'm used to SIMD operations where it's quite usual that you have latency = 1 and reciprocal throughput = 3. This is a narrow path to walk and not one where "tiny bit" of extra overheads will be welcome. I guess we'll see - I would like RISC-V to succeed, but worry that all these resizable models will be quite slow for the codes I write now.
- imtringued 8y agoAre there any cores available that implement it?