3 ms·
I suspect the reason the author is seeing very shallow trees for Nvidia might be because the lower levels are done fully behind the scenes: https://forums.deve
by pixelesque 4y ago
I suspect the reason the author is seeing very shallow trees for Nvidia might be because the lower levels are done fully behind the scenes:
https://forums.developer.nvidia.com/t/extracting-bvh-from-optix-to-manually-traverse/235250 https://forums.developer.nvidia.com/t/extracting-bvh-from-op...
As someone who deals with BVHs a lot for ray intersection, I find it pretty difficult to believe that leaf nodes with that number of primitives will be anywhere near performant, even with fast dedicated hardware like the RT cores.
It's true that the Nvidia cards have better intersection performance than ray/box tests, but I don't believe it's in the 100x ratio range which I suspect would be needed if the BVHs were that shallow and leaf nodes that large.
- frogblast 4y agoI strongly suspect the reason Nvidia trees are so shallow is that NSight simply isn't showing the actual tree structure, probably because Nvidia considers that proprietary. It appears to just list all the leafs of a tree in a big flat list. But there definitely is a tree in there.
- Arrath 4y agoI'm very curious to see it unrolled down to its actual structure.
- kevingadd 4y agoPerhaps the rest of it isn't a tree and is some other optimized data structure? Like some sort of spatial hash or sort
- sounds 4y agoSince cache hit ratios are so central to fast GPU code, the tree structure doesn't have to be exotic, likely the secret sauce is how it performs in the caches.
- graffix 4y agoAt work I inherited a raytracer codebase with a severe memory bloat problem on terrains. The size of terrain BLASes is precisely what one would expect from a bog-standard BVH with branch factor 2, so I'm sure you're right. This is on Turing. Nvidia would've been motivated to de-risk the introduction of RTX by making boring choices. You may well see different results on later archs.
- TinkersW 4y agoIsn't wide BVH how embree works, 1 ray vs SIMD width boxes.. maybe Nvidia is simply doing the same thing but with the wider GPU SIMD(32 I believe).
- berkut 4y agoYes, but normally 4- or 8-wide is the norm: the wider you go the more sorting you have to do to traverse things in order or find the nearest hit which has an overhead (hardware may help with this, but it's still an overhead). Previous indications from Nvidia about their BVHs don't seem to show anything about very shallow trees for any of the BVH algorithms that OptiX supports (scroll to bottom for reverse visualisation of a BVH hierarchy on top of the Stanford Bunny model): https://drive.google.com/file/d/1B5fNRFwv2LsGlCBJ8oKYRiiDUtLMR4TY/view https://drive.google.com/file/d/1B5fNRFwv2LsGlCBJ8oKYRiiDUtL...