4 ms·
Great project. Have you looked at the Convex C-Series architecture? They copied the Cray vector idea. Starting with the C2 (I think) there was also a Vector
by kbob 8y ago
Great project.
Have you looked at the Convex C-Series architecture? They copied the Cray vector idea. Starting with the C2 (I think) there was also a Vector Mask (VM) register which described which vector elements had valid data. Since the C-Series (and the Cray, I think) processed vectors serially, vector ops using the VM would only spend time on the valid elements. It was targeted at codes like this:
for (i = 0; i < n; i++) {
if (A[i] < 34) {
D[i] = A[i] * B[i] + C[i];
}
}
That would translate into code like this. (Actual instructions and mnemonics lost to the mists of time; this is pseudo-assembler.)
ld vl, #N ; assume N <= 128 (-:
ld.v v0, A ; load A
ld s0, #32
less.v v0, s0 ; stores boolean vector into VM
ld.v.t v1, B ; load B where mask == true
mul.v.t v3, v1, v0 ; calc A*B under mask, store in v3
ld.v.t v2, C ; load C under mask
add.v.t v3, v3, v2 ; calc A*B+C under mask, store in v3
st.v.t v3, D
I am ignoring the difference between I32, I64, F32 and F64 because I don't remember how those were coded into the mnemonics, sorry.
There were also instructions to load and store the VM to a scalar register pair.
The Mill architecture, if I understood the lectures, only has a masked store instruction; other vector instructions calculate the values for all elements and maintain an "invalid result" bitvector. An exception is only triggered if an invalid result is actually stored.
- mbitsnbites 8y agoThanks for the reference - I have not heard about the Convex C-Series before so I'll be sure to look it up. While the MRISC32-A1 does serial vector processing, the idea is that you should be able to do parallel processing with the same ISA, so I'm not certain that masking in that fashion is a good option for the MRISC32. Also, having a vector mask register like the Cray limits the size of the vector registers (e.g. to 32 elements in the MRISC32, or 64 elements in a 64-bit architecture).
- childintime 8y agoThe Mill is far ahead of everybody else, and even vectorizes regular for loops. A most beautiful architecture, what it needs is an implementation, and get out of their bubble. Intel and the Mill are destined for each other, but they don't seem to know how to join forces.
- gbrown_ 8y ago> but they don't seem to know how to join forces. I don't understand what you mean by this? The Mill team seem intent of producing a chip of their own.
- childintime 8y agoIndeed. Intel sees the Mill as insignificant, and itself as the Gorilla in the room. But in terms of the architecture, the Mill is on the order of 10x better. That's incredible. While Intel tends to rely on strong-arming it is the #1 Mill enemy while it has to go alone. But it should definitely not be so.