5 ms·
Vectorized Production Path Tracing
- ykl 9y agoThis is one of my favorite graphics/rendering papers from 2017; great work from the entire DWA Moonray team! In addition to the main paper, there are some other materials that were presented at HPG and SIGGRAPH 2017 by the same team. HPG Slides: http://www.highperformancegraphics.org/wp-content/uploads/2017/Papers-Session5/HPG2017_VectorizedProductionPathTracing.pdf http://www.highperformancegraphics.org/wp-content/uploads/20... SIGGRAPH Course Notes (the relevant part begins on Page 35): https://jo.dreggn.org/path-tracing-in-production/part1.pdf https://jo.dreggn.org/path-tracing-in-production/part1.pdf SIGGRAPH Slides: https://jo.dreggn.org/path-tracing-in-production/MoonrayV3.pdf https://jo.dreggn.org/path-tracing-in-production/MoonrayV3.p... Also, the actual shading system they used was also presented at SIGGRAPH 2017: http://blog.selfshadow.com/publications/s2017-shading-course/dreamworks/s2017_pbs_dreamworks_notes.pdf http://blog.selfshadow.com/publications/s2017-shading-course...
- pixel_fcker 9y agoIt's interesting work but it does seem like the speedup they get is pretty poor (1.3x to 2.3x) when going from 1 to 16 SIMD lanes. It seems like the overhead from all the queue management and AOSOA transformation must negate most of the benefit of the parallelization? They also mention the fact that programming the system is hard, and plugins must fall back on single-lane code paths until they can be coded into the system proper. I would assume all the ray sorting makes it extremely difficult to use any bidirectional methods as well.
- Jonanin 9y agoWhy aren't animation houses using GPUs for rendering? It seems like they could get another order of magnitude speedup with CUDA or OpenCL.
- dtf 9y agoPrice, I think. Small outfits do use GPUs in production. edit: I think another reason might be something to do with the gigantic asset sizes (eg textures) typical in big studio jobs.
- MrScruff 9y agoThey already own large cpu only render farms. Production render scenes commonly require upwards of 64Gb of memory for texture and geometry caching. It is used in some niches, and there are efforts to leverage both for interactive rendering scenarios.
- svdree 9y agoYou can, if the scene fits into GPU memory. Otherwise the data transfers will quickly dominate your render time.
- Asooka 9y agoI wonder if chips like the new Intel with AMD integrated GPU would change things. You should be able to simply hand off memory to the GPU to do its thing and read the result by just mapping physical pages back into your address space.
- w0utert 9y agoIt's not that simple. The basic idea behind path tracing maps very well to GPU's, so you can get huge speedups there (10x-100x depending on the scene, geometry, representation, etc). For example path tracing signed distance fields (volume textures, basically) can be done extremely efficiently on GPU's. Things get a lot more difficult when the scene becomes more complex, if it needs to be animated, etc though. You need much more advanced forms of visibility detection/object culling, track scene changes, need to have an much bigger working set (models, textures, metadata) in memory, etc. The amount of code that needs to run compared to 'just the path tracer' starts to far outweigh what the GPU can efficiently process. Hybrid solutions are of possible and widespread, but it is not easy to implement those in a way that will not annihilate the speedup you get from the parts running on the GPU because of synchronization, copying memory around between CPU <-> GPU etc. I can imagine that the kinds of rendering pipelines animation studios use likely depend on hundreds of individual tools from different suppliers, so it would be next to impossible to integrate and the full rendering pipeline efficiently if it runs partly on CPU and partly on GPU. But maybe some parts could be?
- gt_ 9y agoGood summary. I don’t think many 3D artists would disagree that ‘Redshift’ is slightly ahead of the small pack of GPU renders available, but switching has repercussions and limitations throughout the production. These CPU renderers will handle anything you throw at it.
- swerner 9y agoAs others mentioned, asset sizes are one problem. I'm currently involved in a production where render nodes got upgraded to 128GB RAM, since 64GB wasn't enough. Also, path tracing may superficially at first glance look like a good fit for the GPU, but once you look closer, it becomes less so. With incoherent rays and surface shaders, memory access quickly becomes the bottleneck and the GPU cores stall. There are approaches to improving that (batching and sorting), but those can add quite some complexity. Out-of-core memory for GPU rendering can allow for scene sizes to exceed VRAM, but then memory access becomes even slower. Again, with sorting and moving things to VRAM on demand there are paths to reduce the impact of it all, but those are again not trivial given the incoherent access pattern of simulating millions random walks. [source: I am a developer working on the Cycles render engine (Blender) for a production company.]
- science404 9y agoBelieve it or not, performance.. Embree holds up against Nvidia GPUs and Optix, at least it did in 2014 (see: https://embree.github.io/data/siggraph2014_embree_paper.pdf https://embree.github.io/data/siggraph2014_embree_paper.pdf)
- chaintip 9y agoWow, this looks nice.
- gt_ 9y agoThe featured image looks incredible.
- deleted 9y ago[deleted]
- meuk 9y agoWouldn't pathtracing (and raytracing) be a better fit for a manycore architecture? (Or, when memory becomes a bottleneck, a distributed architecture)