14 ms·
It's interesting work but it does seem like the speedup they get is pretty poor (1.3x to 2.3x) when going from 1 to 16 SIMD lanes. It seems like the overhead fr
by pixel_fcker 9y ago
It's interesting work but it does seem like the speedup they get is pretty poor (1.3x to 2.3x) when going from 1 to 16 SIMD lanes. It seems like the overhead from all the queue management and AOSOA transformation must negate most of the benefit of the parallelization?
They also mention the fact that programming the system is hard, and plugins must fall back on single-lane code paths until they can be coded into the system proper.
I would assume all the ray sorting makes it extremely difficult to use any bidirectional methods as well.