5 ms·
I'm adapting QOA to the Dreamcast for streaming music. I've modified it to use 32-bit little endian values instead of 64-bit since (besides avoiding having to b
by TapamN 3y ago
I'm adapting QOA to the Dreamcast for streaming music. I've modified it to use 32-bit little endian values instead of 64-bit since (besides avoiding having to byteswap) the CPU in the Dreamcast is badly suited for 64-bit shifts.
For converting a "slice" from the 64-bit format (which is has a 4-bit value in the highest bits, and 20 3-bit values organized from high to low. Each 3-bit value is read from the high to low), I treated it as two separate 32-bit values. The bottom two bits of are combined to get the 4-bit value, and the rest contains two pairs of 10 3-bit values organized low to high.
So this:
sample = (slice >> 57) & 0x7; //Get sample
slice <<= 3; //Advance to next sample (64-bit shift)
Became this:
sample = slice & 0x7; //Get sample
slice >>= 3; //Advance to next sample (32-bit shift)
I think the conversion to 32-bit little endian was something like an 8-10% speed up, with no other changes? (This was a few months ago, and I didn't write down what the exact difference was.) The overhead big endian conversion and 64-bit shifts would become even more noticeable after further optimizations. (Replacing qoa_lms_predict with inline asm to use an integer multiply-accumulate instruction that GCC can't generate got something like another 20-25% speedup. Properly written asm would be even faster.)