5 ms·
Lately there has also been an effort by a hobbyist to use an MPEG1 decoder (PL_MPEG)[1] on the N64[2]. Disclaimer: I wrote PL_MPEG, but not the N64 port. [1]
by phoboslab 7y ago
Lately there has also been an effort by a hobbyist to use an MPEG1 decoder (PL_MPEG)[1] on the N64[2].
Disclaimer: I wrote PL_MPEG, but not the N64 port.
[1] https://github.com/phoboslab/pl_mpeg https://github.com/phoboslab/pl_mpeg
[2] https://www.reddit.com/r/n64/comments/dr15py/i_just_started_a_dragons_lair_port_to_n64/ https://www.reddit.com/r/n64/comments/dr15py/i_just_started_...
- giovannibajo1 7y agoHi, I'm the guy working on that. I've actually since moved to port a H264 implementation to N64. It's been a long journey and I'm now at around 18 FPS, after nights of manual RSP assembly optimizations, vectorizing most of the intra-prediction and inter-prediction algorithms. I want to reach 30FPS so there's still some work to do.
- fulafel 7y agoWhat kind of vector operations does the N64 support for his case?
- giovannibajo1 7y agoFor H264? Basically everything. Vector registers are 8 lanes, signed 16-bit, so they map quite well to per-pixel calculations on each plane (YUV), which is what video codecs do, as you can process 8 pixels at a time, and you have 16-bit precision to handle intermediate results. The most complex hurdle is that RSP only has 4K of RAM so you need to DMA in and out macroblocks a lot (especially since I can't possibly rewrite a FULL h264 decoder in RSP assembly, not in this lifetime: I need to write only specific performance-sensitive algorithms, while the bulk of the decoder stays in C; this means that the same data ends up going in & out the RSP a lot, especially since the H264 decoder I'm using is not aware of this problem). This said, RSP DMA is even rectangle based, so it's another perfect fit: I can DMA a macroblock by specifying the pointer in RAM, width and height (usually 16x16, but some algos works on sub-partitions of 8x8 or 4x4) and the stride (screen width), so that a single DMA call will transfer the block from the middle a frame, skipping the rest of the data. Vector multiplications in RSP were designed to write DSP-like filters, so they map quite well to the pixel filters required by H264. There are several different multiplication instructions for different fixed point precisions, and there's even one that automatically adds 0.5 (in the correct fixed point precision) which is also a common pattern in FIR filters, and also used in H264. Saturation (VCH/VGE/VLT opcodes) is also supported; this is useful as most algorithms eventually need to saturate the calculated value in the 0-255 range, so that's another thing which usually require 1 clock cycle for 8 pixels. When working with 4x4 partitions, half of the vector lanes are ignored; when writing back to memory, you need to do a read / combine / write sequence (as you may want to write 4 pixels and keep the existing 4 pixels, but vector writes will write 8 pixels); in this case, the VMRG instruction is used, which basically allow to combine two vector registers into one, with a bitmask to specific where to get each lane frame. For IDCT, it comes very handy that most RSP opcodes allows to do partial broadcasts of the lanes of one of the input registers; this allows to keep a 4x4 matrix into 2 consecutive registers and then play some tricks with broadcast to multiply by rows and by columns (which is required by IDCT where you need to compute A' x B x A, with A&B being 4x4 matrices, so if you expand that you will see that you need to rotate vectors a lot). So well, it's actually a pretty good fit. PS: in the Gamasutra article, it shows the RSP code used to do colorspace conversion (YUV->RGB). The article says that it give a big boost (and I can believe it: especially in MPEG1, CSC is like 30% of decoding time), but I brought it to basically 0% by letting the RDP do it (RDP is the GPU in N64). In fact, the RDP supports YUV textures: so in my H264 player, the RSP just does the interleaving (that is, merges the 3 separate Y, U, V planes into one) and then asks the RDP to blit a textured rectangle in the correct format. The RDP even runs in parallel to both RSP and CPU. It might be that, back in 2000, this wasn't fully documented by Nintendo, though I found several references in old Nintendo docs. I can't see otherwise why it wasn't used. Once you reverse engineer how to pass the correct constants, it works really well and brings the CSC cost to basically zero.
- fulafel 7y agoReally fascinating. The dma latency must be pretty small? The 4k of ram and DMA makes the programming model a lot like the Cell, I wonder if experiences with game evs using the RSP microcode encouraged the Cell design. I also wonder how much in common this HW has with the SGI GPUs of the day...
- giovannibajo1 7y agoThe DMA transfers 64-bit words per each bus clock cycle between the main shared memory (RDRAM) and the internal RSP 4K DMEM (or IMEM, to transfer code). So it's quite fast, but you need to remember that the main RDRAM is shared among the main CPU and the whole RCP (eg: it's also used as video memory for textures and frame buffers by the RDP), so contention is really high.
- phoboslab 7y agoThanks for the write-up! This is really interesting. I wouldn't have thought it's possible to get H264 running at reasonable speeds on the N64. Congratz for pulling it off!