6 ms·
Moana Motunui Renderer on GPU
- skunkworker 6y agoFrankly that is incredible that they were able to get the scene to render in only 5 hours per frame on just a single 2070 with 8gb VRam
- iamgopal 6y agoWhole T2 must have rendered for that much processing power.
- corysama 6y agoThe References at the bottom are all also excellent articles. The "Swallowing/Digesting the Elephant" pair are about simply getting the scene to load.
- andrew_ 6y agoI just have to say that this is really excellent work. Especially impressed as I've watched this movie at least 100 times over the last year. My 20 month old daughter is infatuated with it (not to worry, we usually limit viewing to a song-scene or two per sit down). The songs are involuntarily memorized and cries of "mana song" (sic) can be heard each time the toddler gets into the car. To the point that my wife, 3 month old daughter, the oldest, and myself are dressing up for Halloween as TeFiti, kakamora, Moana, and Maui, respectively. If the OP should see this - what prompted you to want to attempt it, was it just the availability of the dataset or were you drawn into the story by screaming children as well?
- chellmuth 6y agoI thought it would be a fun challenge. It is an amazing soundtrack though, I think your daughter has great musical taste :)
- timClicks 6y agoIn case anyone was wondering, "Moana" means sea/lake and "Motunui" means "big island" in many languages of the Pacific.
- LukeShu 6y agoTo answer the implied question at a different level: The scene being rendered is the fictitious island Motunui from the 2016 Disney film Moana.
- dahart 6y agoThis out of core pipeline is a super cool idea! I work on OptiX and this is the first time I’ve seen someone try this kind of thing. There are different kinds of out of core rendering, and not all of them would be guaranteed to finish without running out of memory. This one is, as long as all the individual meshes can fit in GPU memory. Chris are you on HN? I am curious whether some savings might be possible here by not snapshotting the IASes and GASes. Instead you can snapshot only the raw geometry, and then rebuild the GASes and IASes before every launch. It may be faster to rebuild the BVHs than to copy them from CPU ram! *Oh incidentally looking through the code I also noticed that the GASes could be compacted which will save a lot of memory and time/bandwidth on the snapshot copies during rendering. Since you do compact the IAS, I am assuming it's problematic to try and compact the GASes? In any case, nice work!
- chellmuth 6y agoThanks! I really enjoyed working with OptiX. BVH re-building potentially being faster than the associated transfer cost is something that I didn't consider, I'll have to try that out and compare. As for the lack of GAS compaction, I added IAS compaction to squeeze all of the beach's debris into a single snapshot, and just never got to GAS compaction. Right now each rendering pass is only computing one sample per pixel- so in addition to reducing the BVH footprints, I'm hoping I can also hide the transfer costs by computing many more samples per pass.
- dahart 6y agoYes if you have a viable way to compute multiple samples per pass and store the updates, you’ll definitely be able to amortize the cost of the snapshot copies and save a lot of time. It’s been a while since I played with Moana’s meshes, so I don’t remember what GAS compaction will buy you exactly, but generally speaking it’s common to see the compacted GAS end up at about 50% of the size of the first build. If you’re allocating and building multiple GASes in parallel, then it might mean you have to defrag your snapshot in a separate pass using GAS relocation. Once you put all that together, it is possible that copying is faster than rebuilding, so just be aware I’m not certain that my suggestion to build every time is a winner in your case. I can also envision some other strategies that may save considerably on snapshot copies in many cases ( and may be difficult to implement :) ). If you keep going with this and would like more support or suggestions or perf tips, please feel free to get in touch via the OptiX forum (Feel free to DM me there if you like). We’d be happy to try to help you squeeze out more samples per second.
- CamperBob2 6y agoInteresting. The CPU-based texture lookups sound murderous, though. Why not convert the PTEX textures to a format the GPU can handle?
- TomVDB 6y agoHere's what I just noticed in a paper that's was linked at the bottom of this article. https://arxiv.org/pdf/2001.02620.pdf https://arxiv.org/pdf/2001.02620.pdf: "Although OSPRay already supported textures, it only supported image textures, however Ptex is a geometry based texture format baked on top of the underlying meshes [Burley and Lacewell 2008]. Not only does this mean there is no reasonable way these textures could be converted to 2D images for use in OSPRay, but that OSPRay’s entire view of how textures can be applied to geometry—which was inherently based on image textures—would have to change."
- dahart 6y agoNot the author, and it's a totally reasonable suggestion, but having done it, one decent reason is that baking the Ptex into textures that actually fit in the 2070's ram requires making camera specific resolution decisions, adding gutters between adjacent textures, and probably still using out of core texture loading anyway. In other words, it's doable but not easy. In the Moana dataset, the textures consume quite a bit more space than the geometry IIRC.
- ahupp 6y agoThis is really cool work. What kind of speedup does the GPU give for this kind of workload, e.g how long does a CPU-only implementation take?
- packetslave 6y agoFor context, PBRT (a popular open-source CPU renderer) took a crack at it: Total rendering time using a 12 core / 24 thread Google Compute Engine instance running at 2 GHz with the latest version of pbrt-v3 was 1h 44m 45s.
- dahart 6y agoYeah but how long does PBRT take if limited to 8GB ram? ;)
- packetslave 6y agoI mean... "When rendering the latest Moana island scene, pbrt-v3 uses 81 GB of RAM to store the scene description. Today’s pbrt-next uses 41 GB"
- dahart 6y agoOh that’s cool, I didn’t realize pbrt-next was that much leaner. Yeah I was asking a nonsense question, but just tongue-in-cheek trying to hint that we shouldn’t compare PBRT run times to Chris’ project to render Moana on a 2070, they’re solving different problems. Chris’ renderer is also brand new and hasn’t had a chance to evolve the way PBRT did. But all that said, it’s very impressive IMO to render the whole scene with (cpu)Ptex using such a small GPU without it taking far longer. I happen to know for a fact that it’s possible to render Moana much faster than cpu-pbrt if you have access to one of the 48 GB models and don’t need to use out-of-core techniques.
- chellmuth 6y agoI tried to run the pbrt scene locally for timing comparisons, and it kept running out of memory. Thanks for that quote- I realize now that I forgot to switch to the pbrt-next branch. In general, this project has a ton of room left for optimization, so I hope to catch up and maybe pass pbrt. But I'd imagine that a commercial cpu renderer would also perform much better than pbrt.
- TomVDB 6y agoIf I understand this correctly, he renders the whole scene in batches to reduce the memory requirements. Since this is path tracing, how does this work in terms of instances from one batch influencing the lighting of a different batch?
- dahart 6y ago(Also if I understand correctly) there are two kinds of batches here - the instances and the path segments (or path "depth"). Path segment batches are the outer loop, and instance batches are the inner loop. So in terms of path tracing, the instances being batched is a detail that is hidden from the rendering algorithm. Each launch (or batch of instances) performs partial updates to the render buffers (tmax, normal, etc.), until all the instances have been processed, and then the tracing of the next path segment begins, which loops through the instance batches again.
- mrb 6y agoI was comparing the reference images with the rendered ones, and this tree is not at the same location: https://imgur.com/a/JIpgAb8 https://imgur.com/a/JIpgAb8 Some low vegetation (grass) under the palm tree on the left is also absent. Is this part of the limitations the author discloses?: «Other features of the scene are possible to render but out of my initial scope, notably subdivision surfaces and their displacement maps, and a full Disney BSDF implementation»
- asdfasgasdgasdg 6y agoIt may also be a simple divergence in the data used to render the scenes. The article talks about some material colors that aren't present in the data that make pixel perfect rendering impossible. It wouldn't surprise me if they added or removed a tree or two between whenever the data was exported and when the example renders were made.
- etaioinshrdlu 6y agoIs this work sponsored by anyone, or for a competition, or is it just for fun?
- chellmuth 6y agoJust for fun
- VikingCoder 6y agoI've always wished Pixar would do what Id software did, releasing an Open Source version of some small portion of their movies for analysis, like this Moana scene. Because the question on my mind is, given the improvements in rendering, there's some year in which we can now render in real time, what took a long time for Pixar to render. So, what is it? I believe I read that creating the 3D version of Toy Story ran at 24 fps, averaged across the entire film. So, basically, we already crossed the threshold for Toy Story. (But maybe that was only if you have a rendering farm?) So, where are we at? Could we render a plausible Toy Story in real time now? Bug's Life? Monsters Inc?
- thdrdt 6y agoAll Blender movies can be downloaded for this reason: https://cloud.blender.org/open-projects https://cloud.blender.org/open-projects Last month they discovered a bug in Blender that was the cause of very slow hair rendering. The fix dropped render times for some scenes from 4 minutes to 40 seconds. So it can indeed be useful to check if things can be improved for old movies. But I am not sure realtime is an option right now. Most movies also use post processing per frame wich also takes some time.
- jsheard 6y ago> So, where are we at? Could we render a plausible Toy Story in real time now? Bug's Life? Monsters Inc? The Kingdom Hearts games have areas based on various Disney IPs, and the last one came out fairly recently so it makes for a nice comparison Pixar Toy Story vs Realtime Toy Story: https://www.youtube.com/watch?v=tkDadVrBr1Y https://www.youtube.com/watch?v=tkDadVrBr1Y Disney Frozen vs Realtime Frozen: https://www.youtube.com/watch?v=q8vCQqg6SYg https://www.youtube.com/watch?v=q8vCQqg6SYg
- aidenn0 6y agoKH3 from 3 years ago clearly is missing dynamic shadows (e.g. when Woody moves his head, the lighting around his face changes based on shadows cast by his hat and nose). Also the original toystory appeared to have at least 2 surfaces for reflecting off of the floor while I didn't see any reflections off of the floor in KH3. Granted that's 3 years old now. Also, TS was definitely rendered higher than 1080p to make the transfer to 35mm film for large screens.
- fxtentacle 6y agoMost of the comments in here seem to be missing the most important aspect. This is open source: https://github.com/chellmuth/gpu-motunui/ https://github.com/chellmuth/gpu-motunui/ Let me explain. GPU rendering of cinema quality scenes by itself is not that new. V-Ray GPU was around in 2013 and already had an impressive showreel back then: https://www.youtube.com/watch?v=RYPFY5OUzdk https://www.youtube.com/watch?v=RYPFY5OUzdk I'd count Octane and Redshift as the next big contenders, who together with V-Ray switching from perpetual to subscription pricing, took over the market. Here's an example of a Redshift render from 2016: https://www.youtube.com/watch?v=to8yh83jlXg https://www.youtube.com/watch?v=to8yh83jlXg Lately, Pixar has been handing out Renderman GPU betas, which has already been used in the new jungle book movie: https://youtu.be/tiWr5aqDeck?t=171 https://youtu.be/tiWr5aqDeck?t=171 Along with this development, the price has gone down. From $1200 per V-Ray license, to $600 per Redshift license and now it seems like the previously prohibitively expensive Pixar renderer will soon be offered for $500. And recently, people have been using Houdini ($6000 per seat) to export their scenes as USD (free open file format) so that they can use Blender 2.81 for rendering (also free). But the conversion to get things into Blender is work. This renderer can use the Disney production data directly, which is very convenient if you want to drop it in at the last minute for cost saving. GPU rendering has gone from novelty in 2013 to established in 2016 to become the new default in 2018/2019. And this released as open source to me implies that the commercial renderer market will soon be dead. It looks "good enough" for advertisements and archviz, which is where the money is made.
- alkonaut 6y agoWhat is the performance difference between state of the art CPU vs. GPU? e.g. Embree vs. OptiX?
- dahart 6y agoFor ideal scenes that fit completely in GPU memory and use hardware meshes and hardware texture filtering and GPU shading, rendering on a GPU is often in the range of 10x to 100x faster than CPU (as reported by the people writing production renderers today). Since production scenes are commonly larger than GPU ram and even CPU ram too, it complicates things. Embree is a renderer and OptiX isn’t, so they can’t be directly compared, but I think it’s fair to say that OptiX based renderers are typically significantly faster for ideal scenes that fit in memory and use hardware supported primitives, while there are cases where Embree has some huge advantages, like scenes that don’t fit in memory, and Embree has oriented bounding boxes that can make some kinds of scene geometry render much faster.
- packetslave 6y agoA couple of other attempts at rendering the Moana scene with something other than Hyperion: - Matt Pharr (author of pbrt) -- 6 part series: https://pharr.org/matt/blog/2018/07/08/moana-island-pbrt-1.html https://pharr.org/matt/blog/2018/07/08/moana-island-pbrt-1.h... - Joe Schutte (WDAS): https://schuttejoe.github.io/post/disneybsdf/ https://schuttejoe.github.io/post/disneybsdf/
- virtualritz 6y ago3Delight (scroll down to the bottom of the page) – https://www.3delight.com/documentation/display/3DLC/Cloud+Rendering+Speed https://www.3delight.com/documentation/display/3DLC/Cloud+Re... Their conversion script for the original Disney asset is available on GiLab: https://gitlab.com/3Delight/moana-to-nsi https://gitlab.com/3Delight/moana-to-nsi
- boulos 6y agoNice work! In addition to dahart’s suggestions, I’d be curious to see some of your profiling that you did. Like, what’s the breakdown of the timing now? Are you PCIe limited, or have you overlapped the I/O with enough compute that it’s only XX% overhead? How much memory do you need for all the geometry and BVHs? (Particularly post compression). I’m curious if this just barely fits on an A100 or similar. Again, nice work!
- chellmuth 6y agoThanks! The current version requires around 18GB of total GPU memory. Right now I am heavily bottlenecked by the memory transfer costs, but there is some low-hanging fruit to amortize them by doing much more compute per snapshot transfer. After that, I'm very curious to see what the profiler will show.
- boulos 6y agoOoh! With Dave’s compression suggestion, I wonder if that gets you below the 16 GB of a T4 (which have RTX).
- moyix 6y agoMatt Pharr recently released an early version of PBRTv4, which has been rewritten extensively to make heavy use of Optix and the GPU. Given that he previously wrote a series on rendering Moana using PBRTv3, I wonder if v4 could be used as a comparison? https://pharr.org/matt/blog/2020/08/19/pbrt-v4-released.html https://pharr.org/matt/blog/2020/08/19/pbrt-v4-released.html
- virtualritz 6y agoThis looks mostly like an academic or 'for fun' exercise. People run Doom on their fridge for fun so why not squeeze the Moana asset through the bottleneck that sits between your CPU and your GPU? :) Is it practical/useful? Let's put the timing in perspective. A commercial CPU production renderer, 3Delight, has timing for the Moana asset rendered at 4k on their website.[1] Time: ~34 minutes. I asked them for details about the settings they used before posting this as the page only lists resolution. 4k resolution, 64 (shading) samples per pixel (spp), ray depths: diffuse 2, specular 2, refraction 4 (or 3, 3, 5, depending how you count ray depth). Machine was a contemporary 24 core server at the end of 2018. Mind you, the image is fine with 64 spp. Spp are hard to compare between renderers because optimizing path tracers is a lot about sampling. One renderer will converge to something useable with 1k samples while another just needs 64. The 3Delight example is rendering all the geometry as subdivision surfaces with displacement (and their own Ptex implementation for texture lookups). Timing comparisons of a different scene with recent 24core desktop AMD CPUs suggest that this asset would render much faster in 2020.[2] The timing shows the issue with GPUs vs CPUs for this kind of assets. 5h for a 1k (!) resolution image with <= 5 bounces and 1024 spp (samples per pixel). That is terrible. Not using the real (subdivision) geometry and not using displacement. I would love to see a breakdown how much of these 5hs is owed to the fact that the data doesn't fit on the device. Using subdivision surfaces and displacement mapping make the amount of geometry grow exponentially. I.e. the out of core handling would predictably take an exponentially larger part of the render time. Looking at the numbers I regularly get to see when counseling VFX companies on their rendering pipelines I don't see GPU offline rendering going anywhere for complex scenes. And even for simpler scenes where GPUs have an advantage still – with the CPUs AMD is putting out recently the gap is becoming very tight and if you do the math you often pay dearly for having an image a few minutes earlier (not even double digit minutes or hours earlier). Regardless of what Nvidia's marketing and some vendors who IMHO wasted years optimizing their renderers for a moving hardware target may want you to believe. Regarding the latter: another point to consider is that you need to spend time working with/around the hardware limitations/bottlenecks of GPUs for this very the "scene doesn't fit on device" use case. Someone writing a CPU renderer can spend that time working on the actual renderer itself. This kind of software takes years to develop. Go figure. Finally, as I expect this to be downvoted because of what I just said: don't take my word for any of the above. Just try it yourself. The Moana Asset can be downloaded at [3]. A script to convert the entire asset and launch a 3Delight render can be had at [4]. The unlimited core version of the renderer can be downloaded for free, after registering with your email, at [5]. It renders with any number of cores your box has but it adds a watermark if no license is available. Or you thy their cloud rendering. You get 1,000 free 24 core server minutes. Which is plenty to run this test. [1] https://www.3delight.com/documentation/display/3DLC/Cloud+Rendering+Speed https://www.3delight.com/documentation/display/3DLC/Cloud+Re... [2] https://www.3delight.com/page/features/2020-10-06-CPUbenchmark https://www.3delight.com/page/features/2020-10-06-CPUbenchma... [3] https://disneyanimation.com/resources/moana-island-scene/ https://disneyanimation.com/resources/moana-island-scene/ [4] https://gitlab.com/3Delight/moana-to-nsi https://gitlab.com/3Delight/moana-to-nsi [5] https://www.3delight.com/download https://www.3delight.com/download