3 ms·
I'm willing to bet that if you just treated each frame as an image, it would result in some weird stuff when you played them as a movie. > penny per frame Whe
by caturopath 3y ago
I'm willing to bet that if you just treated each frame as an image, it would result in some weird stuff when you played them as a movie.
> penny per frame
Where did this come from?
- leetharris 3y agoI do lots of large scale ML work, this was just sort of a random educated "order of magnitude" guess.
- syntaxing 3y agoSeems like a pretty reasonable estimate, if it cost about $2 a hour to rent a decent GPU, that’s 18s per penny which sounds pretty doable to run one frame.
- erwannmillon 3y agoThink inference time was on the order of 4-5seconds per image on a v100, which you can rent for like .80 cents an hour, though you can get way better gpus like a100s for ~1.1 usd/h now. But ofc this is at 64px res in pixel space. If you wanted to do this at high res, you would definitely use a latent diffusion model. The autoencoder is almost free to run, and reduces the dimensionality of high res images significantly, which makes it a lot cheaper to run the autoregressive diffusion model for multiple steps.