4 ms·
Frankly I'm surprised this isn't much higher quality. The hard thing in the transformers era of ML is getting enough data that fits into the next token or maske
by wantsanagent 2y ago
Frankly I'm surprised this isn't much higher quality. The hard thing in the transformers era of ML is getting enough data that fits into the next token or masked language modeling paradigm, however in this case, inbetweening is exactly that task and every hand-drawn animation in history is potential training data.
I'm not surprised that using off the shelf diffusion models or multi-modal transformer models trained primarily on still images would lead to this level of quality, but I am surprised if these results are from models trained specifically for this task on large amounts of animation data.
- yosefk 2y agoThey're indeed not diffusion models, though they are trained on animation data as well as specifically designed for it (the raster papers at least.) I'm very hopeful wrt diffusion, though I'm looking at it and it's far from straightforward. One problem with diffusion and video is that diffusion training is data hungry and video data is big. A lot of approaches you see have some way to tackle this at their core. But also, AI today is like 80s PCs in some sense: both clearly the way of the future and clumsy/goofy, especially when juxtaposed with the triumphalism you tend to hear all around