3 ms·
Pure speculation/dreaming/brainfarting on my part: While I think DLSS will do a major part of the heavy lifting for upscaling, denoising and frame-generation I
by riggsdk 3y ago
Pure speculation/dreaming/brainfarting on my part:
While I think DLSS will do a major part of the heavy lifting for upscaling, denoising and frame-generation I think the brunt of the AI generation will happen in the engines and/or graphics APIs.
I think the full power of this will come gradually where they allow you to tag certain objects to be neurally generated. You can load in a set of specialized models optimized for certain effects and download new ones as people create them (like shaders).
This will most likely start with "noisy" effects like grass, hair, water, smoke, clouds and other particle effects where you will rarely see two identical frames. Because of their noisy nature they will be easier to fake from a perception staindpoint as you don't need perfect replication/reproducibility as you move the camera around. It needs to be temporally stable between frames of course but if you re-render at a later time from the same angle/position you don't expect the exact same result as before.
Then you have materials of surfaces like dirt, mud, sand, soil, sparkly metal, skin, leather and other similar material that are completely geometry/shader drawn today. They could be used more frequently if they end up being faster than regular shaders.
Neural networks then begin complementing all skeletal animation systems. The artist still controls the large-scale movement but lets the neural net fill in blanks so it better suits the specifics of the scene and interactions with the world.
The renderer can draw certain things the old fashioned way and then one or multiple layers of a low resolution "tagged" scene that via the tags point to all the underlying information on how we want certain things to look, like artist/AI drawn reference images of certain characters. Neural nets then does facial animation (eyes, mouth, grimaces) entirely in pixel space with no interaction with the underlying bone system. You just provide a "mask" of the pixels where you want the effect to happen as well as the input appearance information.
A final "styling" neural net can then pass over the entire image applying any final touches to the scene as it upscales it to high-res. For example doing the color-grading to always fit the scene best and maybe turning a somewhat photorealistic scene into a slight cartoon'esque like style or to fit a given pre-trained model on some sample premade 3D renders/movies (like that example where it applied a realistic look to a GTA-like input scene based on training data filmed while driving around a city or taken from googlestreetmaps)
All steps of the pipeline will be fully configurable to get the visual style and animation you want.
Modders of games will replace certain neural models with better ones just like people already replace shaders and textures.
Question is just, how will this be standardized? How will Vulkan/OpenGL/Metal/D3D provide a "simple" API for using such advanced features?
Graphics cards will change and will come with insane amounts of RAM to fit more and more models into them.
More of the GPU silicon will be dedicated to run neural effects efficiently.
They will quickly take a huge percentage of the GPU RAM than textures and other resources do today.
It will undoubtedly expand to sound generation as well. No more repeated demon growls or bird chirps you've heard a thousand times before.
- ksec 3y agoI may be the minority. But part of me being a nerds absolutely love this crazy idea, and it will brings huge efficiency improvement, cost reduction and potentially quality increase in AAA games. May be all at the same time. Exciting. But another part of me for reason I cant quite put into words right now think this is VERY scary.