2 ms·
> While each lane isn't truly a thread because it doesn't have its own PC the programming model definitely tries to make it seem that way. The threads can termi
by pezezin 2mo ago
> While each lane isn't truly a thread because it doesn't have its own PC the programming model definitely tries to make it seem that way. The threads can terminate at different points too. And again, the ISA isn't a vector ISA. Your register values are scalar.
This is not correct. If you check AMD's documentation there are explicit mentions of vector registers (VGPR), vector ALUs, and vector instructions. The introduction to Chapter 2 describes it as a vector ISA.
> RDNA4 shader programs (kernels) are programs executed by the shader processor. Conceptually, the shader program is executed independently on every work-item, but in reality the processor groups up to 32 or 64 work-items into a wave, that executes the shader program on all 32 or 64 work-items in one pass ("wave32" or
"wave64").
Sources:
https://gpuopen.com/amd-gpu-architecture-programming-documentation/ https://gpuopen.com/amd-gpu-architecture-programming-documen...
https://docs.amd.com/v/u/en-US/rdna4-instruction-set-architecture https://docs.amd.com/v/u/en-US/rdna4-instruction-set-archite...
- MindSpunk 2mo agoA VGPR is not the same thing as a vector register like in SSE4 or AVX. Each addressed register contains a single 32-bit value. A VGPR differs from an SGPR in that each thread in a thread group can have a different value in that register. An SGPR will have a uniform value shared with all threads in a group. An add instruction on an AMD GPU adds two scalar values. If they're in a VGPR then each thread will add two values unique to that thread. A SIMD ISA as is common on a CPU is different because an add instruction explicitly adds a vector of values. xmm1 stores 128-bits of data. VGPR[1] stores 32-bits of data vectored over 32-64 threads in a thread group. Without special instructions a thread can't access the VGPR values stored in other threads.
- pezezin 2mo agoBut those instructions exist, see Section 7.9 "Cross-Lane and Data Parallel Processing (DPP)" AMD documentation describes RDNA as a vector ISA, so I don't understand why you say it is not.