3 ms·
What's next? Dreamfusion Video = Imagen Video (this) + Dreamfusion (https://dreamfusion3d.github.io/ https://dreamfusion3d.github.io/) Fundamentally, I think w
by fzysingularity 4y ago
What's next?
Dreamfusion Video = Imagen Video (this) + Dreamfusion (https://dreamfusion3d.github.io/ https://dreamfusion3d.github.io/)
Fundamentally, I think we have all the pieces based on this work and Dreamfusion to make it work. From the looks of it, there's a lot of SSR (spatial SR) and TSR (temporal SR) going on at multiple levels to upsample (spatially) and smoothen (temporally) images that won't be needed for NERFs.
What's impressive is the ability to leverage billion-scale image-text pairs for training a base model that can be used to super-resolve over space and time. And that they're not wastefully training video models from scratch, and instead separately training TSR, SSR models for turning the diffused images to video.
- stillsut 4y agoI think 3D is going to be really important because I see Generative-AI as the killer app for VR/AR. As it stands, it's very difficult to invest the budget for a dev studio (dozens of high skill people) to build a "VR movie" when the format is so unknown and unpopular. But with generative AI, an indie dev could create their own professionally produced virtual world movie. It's these creatives and risk takers that will find what types of things VR needs to become more popular.
- bredren 4y agoThis direction will provide the visuals but what also must be brought in is a language model and text to speech (TTS) so that you may talk and interact with these things.