6 ms·
For anyone that is curious, this is the paper that describes the core vectorized path tracing architecture in Moonray: http://www.tabellion.org/et/paper17/Moon
by ykl 4y ago
For anyone that is curious, this is the paper that describes the core vectorized path tracing architecture in Moonray:
http://www.tabellion.org/et/paper17/MoonRay.pdf http://www.tabellion.org/et/paper17/MoonRay.pdf
Extracting sufficient coherency from path tracing in order to be able to get good SIMD utilization is a surprisingly difficult problem that much research effort has been poured into, and Moonray has a really interesting solution!
- bhouston 4y agoIs there any comparisons to GPU-accelerated rendering? It seems most people are going that direction rather than trying to optimize for CPUs these days, especially via AVX instructions.
- jsheard 4y agoCPUs are still king at the scale Dreamworks/Pixar/etc operate at, GPUs are faster up to a point but they hit a wall in extremely large and complex scenes. They just don't have enough VRAM, or the work is too divergent and batches too small to keep all the threads busy. In recent years the high-end renderers (including MoonRay) have started supporting GPU rendering alongside their traditional CPU modes, but the GPU mode is meant for smaller scale work like an artist iterating on a single asset, and then for larger tasks and final frame rendering it's still off to the CPU farm. Pixar did a presentation on bringing GPU rendering to Renderman, which goes over some of the challenges: https://www.youtube.com/watch?v=tiWr5aqDeck https://www.youtube.com/watch?v=tiWr5aqDeck
- pixelpoet 4y agoWhat's your opinion on renderers such as Redshift which explicitly target production rendering and support out of core rendering on GPUs? See e.g. https://www.maxon.net/en/redshift/features?categories=631816 https://www.maxon.net/en/redshift/features?categories=631816 (Disclosure: I work on this.)
- berkut 4y agoThose are generally being used on much smaller productions, or at least "simpler" fidelity things (i.e. non-photo CG animation like Blizzards's Overwatch). So for Pixar/Dreamworks style things (look great, but not photo-real) they're useable and provide a definite benefit in terms of iteration time for lookdev artists and lighters, but it's not there yet in terms of high-end rendering at scale.
- virtualritz 4y agoAs someone said above: GPUs are fine & faster as long as your scene stays simple. As soon as you hit a certain scene complexity ceiling, they become much slower that CPU renderers. I would also argue that for this specific task, i.e. offline rendering such frames, the engineering overhead to make stuff work on GPUs is better spent making stuff faster and scale more efficiently on CPUs.[1] I worked in blockbuster VFX for 15 years. It's been a while but I have network of people in that industry, many working on these renderers. The above is kinda the consensus whenever I talk to them. [1] With the aforementioned caveat: if the stuff you work on is always under that complexity ceiling targeting GPUs can certainly make sense.
- bhouston 4y agoSo we just need GPUs with 128GB of ram then? Or move towards the Apple M-series design where CPU+GPU both have insanely fast access to all ram...
- jsheard 4y agoIt's easier said than done, there's consistently a huge gulf between CPU and GPU memory limits. Even run-of-the-mill consumer desktops can run 128GB of RAM, which exceeds even the highest end professional GPUs VRAM, and the sky is the limit with workstation and server platforms. AMD EPYC can support 2TB of memory!
- jb1991 4y agoIt’s not just about memory. The path tracing algorithm is a natural fit for CPU threads but very difficult to design for efficient use of GPU threads. It’s very easy to leave many of your GPUs threads idle due to the divergence, or overflowing the registers, and any number of other things that are very natural to path tracing.
- polishdude20 4y agoSo the idea is one CPU can have hundreds of gigabytes of ram at a time and the speed of the cpu is no problem because you can scale the process over as many CPUs as you want?
- kranke155 4y agoFundamentally yes. Big Studios are CPU Farms. Small Studios and Indie Artists like myself a Lot of us have moved to GPU.
- polishdude20 4y agoIs that because work is usually divided by frame? And usually a frame for these big movies uses more than typical GPU VRAM?
- kranke155 4y agoThe work is always divided per frame yes. Everything is baked so that happens easily even with water simulations, they’ve been reduced down to some form of geometry or something similar that can be rendered frame by frame. The VRAM is indeed one of the main issues. But as someone else I believe cost per final pixel is still lower in CPU. That was particularly true during the GPU shortage.
- chris37879 4y agoIts more to do with the movies using higher fidelity assets than you'd typically use for a game, which is what GPUs are made for. In a movie, a single element of a scene might have as many polygons as an entire character in a video game, and its because of the differences in how you 'film' them. Imagine a brick wall rendered for a video game vs one rendered for a movie. The one for the game is probably going to be a plane with a couple textures on it because the wall is something that's a background element that the player isn't going to be up close with. Whereas in the movie, the wall is more likely to be made of a handful of individually modeled bricks with much more detailed surface textures because maybe the director wants the camera to be really close to the surface of that wall and pull out to a wider shot, so that means the individual brick you start zoomed in on might have a 4k texture for itself alone, whereas the entire wall in the video game could easily be a single 4k texture since the player doesn't get close enough to notice the missing detail. Now multiply that level of detail across every rendered thing in the scene, because the director may want to reframe the shot, or you need realistic lighting to sell that a rendered thing is integrated with filmed footage and 'real'. Every little bit of that detail adds more data you have to track. So in my wall example, you might have 2 or 3 4k textures vs literally hundreds for all the bricks, grout, defects, chipped faces, etc of a movie quality wall.
- aprdm 4y agoI think it's mostly a question of price currently. AMD CPUs are much cheaper per pixel produced that GPUs
- johnvanommen 4y agoI talked to some folks who worked there, years ago, and was surprised they didn't use GPUs. I got the impression that the software was largely based on code that dated back to the 1990s.
- softfalcon 4y agoThis is exactly why I jumped into the comments. I was hoping someone had some relevant implementation details that isn't just a massive GitHub repo (which is still awesome, but hard to digest in one sitting). Thank you!
- bri3d 4y agoThis paper is the first place I've found a production use of Knights' Landing / Xeon Phi, the Intel massively-multicore Atom-with-AVX512 accelerator system, outside of HPC / science use cases. And for this use case, it makes perfect sense!
- ace2358 4y agoIs that sort of code able to ‘just run’ on their arc series of GPUs? It feels like intel kinda hates AVX512 on the cpu side (or wants to upsell you for it) so I’m wondering if they turned those cards into their GPUs
- zokier 4y agoI don't think Arc and Xeon Phi have much commonality at all
- mhh__ 4y agoProbably not. Even if the code was entirely reliant on the compiler performing vectorization (i.e. GPU & CPU instruction sets are at very the least nominally very different), GPU programming is quite different even to "lots of little CPUs"-style.
- Pet_Ant 4y ago> “lots of little CPUS”-style Known as “manycore”: https://en.m.wikipedia.org/wiki/Manycore_processor https://en.m.wikipedia.org/wiki/Manycore_processor
- bri3d 4y agoIt seems like Dreamworks had to move acceleration to GPU using CUDA and OptiX instead of SIMD/AVX and Embree - the released OpenMoonRay code supports both. It kind of feels like Dreamworks got burned a little by Intel here - they invested a ton of research effort (alongside Intel) in Embree, SIMD, and Knights' Landing, and then had to add another implementation for OptiX/GPU/SIMT anyway.
- aprdm 4y ago
- pengaru 4y ago> Extracting sufficient coherency from path tracing in order to be able to get good SIMD utilization is a surprisingly difficult problem Huh, I'd have assumed SIMD would just be exploited to improve quality without a perf hit, by turning individual paths into ever so slightly dispersed path-packets likely to still intersect the same objects. More samples per path traced...
- berkut 4y agoIf you only ever-so-slightly perturb paths, you generally don't get anywhere near as much of a benefit from monte carlo integration, especially for things like light transport at a global non-local scale (might plausibly be useful for splitting for scattering bounces or something in some cases). So it's often worth paying the penalty of having to sort rays/hitpoints into batches to intersect/process them more homogeneously, at least in terms of noise variance reduction per progression. But very much depends on overall architecture and what you're trying to achieve (i.e. interactive rendering, or batch rendering might also lead to different solutions, like time to first useful pixel or time to final pixel).
- xeromal 4y ago[flagged]