5 ms·
Well, not many DC emulators will run my code anymore, and eventually none will. :P On the DC, I used to use serial port with a USB adapter that manages 384k bau
by TapamN 8y ago
Well, not many DC emulators will run my code anymore, and eventually none will. :P On the DC, I used to use serial port with a USB adapter that manages 384k baud, but upgraded to a BBA several years ago.
I really, really, don't like the idea of doing a bunch of development on an emulator, then finding out it doesn't work right on real hardware (either not at all or with unexpectedly terrible performance) and having to figure out what all is going wrong. Doing it all on a real console means you find out immediately where any problems arise.
The Dreamcast tries to pretend to be like a PC, but treating it like one will seriously limit performance. The SH-4's terrible cache needs to be babied to get the most out of it, and there's a lot of unusual things you can do with the 3D hardware since it works very differently than other GPUs.
You can definitely do assembly on the DC, and it's really important for performance critical code. GCC can't use the SH-4's SIMD instructions effectively, so you have to use at least inline asm or true assembly. I'm not totally sure what you mean by "in very logical ways". I guess optimal SH-4 code doesn't look logical! Rendering-type code is harder on the SH-4 than the SH-2/3 because it's superscalar and FPU instructions have longer latency than ALU instructions. You have to spend more time figuring out how to hide latencies, and software pipelining becomes more important for performance. Optimal SH-4 assembler code ends up much messer looking than optimal SH-2/3 code.
For example, normalizing an array of 3D vectors in C, using inline assembly to use the inner product (FIPR) and square root reciprocal (FSRRA) instructions, you'd be lucky to get around 24-28 cycles per vector (it'd be even worse if you used pure vanilla C). But with software pipelined assembly, I can manage 8 cycles per vector.
But it's much more complex, not just because it's in assembler. The optimized main loop juggles normalizing 4 vectors at the same time (The loop is 32 cycles, with 4 vectors it averages 8 cycles per vector), and there's code to prepare the loop and exit the loop, and special case code for if there are fewer than 4 vectors to transform. The resulting machine code for the entire asm version is probably around 10 times the simple version.
Another example would be transforming a bunch of 4D vectors by a 4x4 matrix. C with inline asm will take about 16 cycles per vector, but optimized pure asm can do 4 cycles.
A 3D rendering engine I'm working on for the DC does software vertex shaders by basically copy & pasting optimized assembly loops into a single function and adding a bit of glue code between them.
- mkesper 8y agoIs your code inspected by emulator writers? I guess there would be a lot to learn.