4 ms·
With an SPU's 256K local memory and DMA, the ideal way to use the SPU was to split the local memory into 6 sections: code, local variables, DMA in, input, outpu
by corysama 2y ago
With an SPU's 256K local memory and DMA, the ideal way to use the SPU was to split the local memory into 6 sections: code, local variables, DMA in, input, output, DMA out. That way you could have async DMA in parallel in both directions while you transform your inputs to your outputs. That meant your working space was even smaller...
Async DMA is important because the latency of a DMA operation is 500 cycles! But, then you remember that the latency of the CPU missing cache is also 500 cycles... And, gameplay code misses cache like it was a childhood pet. So, in theory you just need to relax and get it working any way possible and it will still be a huge win. Some people even implemented pointer wrappers with software-managed caches.
500 cycles sounds like a lot. But, remember that the PS2 ran at 300MHz (and had a 50 cycle mem latency) while the PS3 and 360 both ran at 3.2Ghz (and both had a mem latency of 500 cycles). Both systems pushed the clock rate much higher than PCs at the time. But, to do so, "niceties" like out-of-order execution were sacrificed. A fixed ping-pong hyperthreading should be good enough to cover up half of the stall latency, right?
Unfortunately, for most games the SPUs ended up needing to be devoted full time to making up for the weakness of the GPU (pretty much a GeForce 7600 GT). Full screen post processing was an obvious target. But, also the vertex shaders of the GPU needed a lot of CPU work to set them up. Moving that work to the SPUs freed up a lot of time for the gameplay code.
- bri3d 2y agoI think one thing that the linked article (which I think is great and I generally agree with!) misses is that libraries and abstraction can patch over the lack of composability created by heterogeneous systems. We see it everywhere - AI/ML libraries abstracting over some combination of TPU, vector processing, and GPU cores being one obvious modern place. This happened on the PS3, too, later in its life: Sony released PlayStation Edge and middleware/engine vendors increasingly learned how to use SPU to patch over RSX being slow. At this point developers stopped needing to care so much about the composability issues introduced by heterogeneous computing, since they could use the SPUs as another function processor to offload, for example, geometry processing, without caring about the implementation details so much.
- 01HNNWZ0MV43FF 2y agoI'm surprised the SPUs were used for post-processing, cause whenever I try to do software rendering I get bottlenecked on fill rate quickly. I believe you, because I've seen it attested in many places, but I'm surprised by it.
- corysama 2y agoThe 1:1 straight-line behavior of fullscreen post processing is much easier to prefetch than triangle rasterization. And, in this case the SPUs and GPU used the same memory. So, no bandwidth advantage to the GPU. The best the GPU could do would be hiding latency better.
- masklinn 1y ago> Both systems pushed the clock rate much higher than PCs at the time. Intel reached 3.2GHz on a production part in June 2003, with the P4 HT 3.2 SL792. At the time the 360 and PS3 were released, Intel's highest clocked part was the P4 EE SL7Z4 at 3.73.
- rasz 1y agoNot to mention both Intel 30 stages deep pipeline and PPC in-order were empty MHz spend mostly on waiting for cache misses.