5 ms·
The rate of progress in ML this past year has been breath taking. I can’t wait to see what people do with this once controlnet is properly adapted to video. Ge
by valine 3y ago
The rate of progress in ML this past year has been breath taking.
I can’t wait to see what people do with this once controlnet is properly adapted to video. Generating videos from scratch is cool, but the real utility of this will be the temporal consistency. Getting stable video out of stable diffusion typically involves lots of manual post processing to remove flicker.
- Der_Einzige 3y agoControlnet is adapted to video today, the issues are that it's very slow. Haven't you seen the insane quality of videos on civitai?
- valine 3y agoI have seen them, the workflows to create those videos are extremely labor intensive. Control net lets you maintain poses between frames, it doesn’t solve the temporal consistency of small details.
- mattnewton 3y agoPeople use animatediff’s motion module (or other models that have cross frame attention layers). Consistency is close to being solved.
- valine 3y agoHopefully this new model will be a step beyond what you can do with animatediff
- dragonwriter 3y agoTemporal consistency is improving, but “close to being solved” is very optimistic.
- mattnewton 3y agoNo I think we’re actually close. My source is I’m working on this problem and the incredible progress of our tiny 3 person team at drip.art (http://api.drip.art http://api.drip.art) - we can generate a lot of frames that are consistent, and with interpolation between them, smoothly restyle even long videos. Cross-frame attention works for most cases, it just needs to be scaled up. And that’s just for diffusion focused approaches like ours. There are probably other techniques from the token flow or nerf family of approaches close to breakout levels of quality, tons of talented researchers working on that too.
- ryukoposting 3y agoThe demo clips on the site are cool, but when you call it a "solved problem," I'd expect to see panning, rotating, and zooming within a cohesive scene with multiple subjects.
- mattnewton 3y agoThanks for checking it out! We’re certainly not done yet, but much of what you ask is possible or will be soon on the modeling side and we need tools to expose that to a sane workflow in traditional video editors.
- Hard_Space 3y agoOnce a video can show a person twisting round, and their belt buckle is the same at the end as it was at the start of the turn, it's solved. VFX pipelines need consistency. TC is a long, long way from being solved, except by hitching it to 3DMMs and SMPL models (and even then, the results are not fabulous yet).
- capableweb 3y ago> Haven't you seen the insane quality of videos on civitai? I have not, so I went to https://civitai.com/ https://civitai.com/ which I guess is what you're talking about? But I cannot find a single video there, just images and models.
- Kevin09210 3y agohttps://www.youtube.com/shorts/ZN-NbdFwfNQ https://www.youtube.com/shorts/ZN-NbdFwfNQ https://www.youtube.com/watch?v=3WWy98ylLT4 https://www.youtube.com/watch?v=3WWy98ylLT4 https://www.youtube.com/shorts/1vqOjYWEF84 https://www.youtube.com/shorts/1vqOjYWEF84 https://www.youtube.com/shorts/jOIb9QbrhZ8 https://www.youtube.com/shorts/jOIb9QbrhZ8 https://www.youtube.com/shorts/C3F_YI84TXA https://www.youtube.com/shorts/C3F_YI84TXA https://www.youtube.com/shorts/4IqJHozY4F0 https://www.youtube.com/shorts/4IqJHozY4F0 https://www.youtube.com/shorts/h3OmBLlm5-g https://www.youtube.com/shorts/h3OmBLlm5-g https://www.youtube.com/shorts/ZT7tuIgSDRk https://www.youtube.com/shorts/ZT7tuIgSDRk https://www.youtube.com/shorts/WnUYbsOMyvs https://www.youtube.com/shorts/WnUYbsOMyvs https://www.youtube.com/shorts/BKKqX2aMlSg https://www.youtube.com/shorts/BKKqX2aMlSg The inconsistencies are what's most interesting in these videos in fact
- capableweb 3y agoNot sure I'd call that "insane quality", more like neat prototypes. I'm excited where things will be in the future, but clearly it has a long way to go.
- dragonwriter 3y agoA small percentage of the images are animations. This id (for obvious reasons) particularly common for images used on the catalog pages for animation-related tools and models, but also its not uncommon for (AnimateDiff-based, mostly) animations to be used to demo the output of other models.
- adventured 3y agohttps://civitai.com/images https://civitai.com/images Go there, in the top right of the content area it has two drop-downs: Most Reactions | Filters Under filters, change the media setting to video. Civitai has a notoriously poor layout for finding/browsing things unfortunately.
- alberth 3y agoWhat was the big “unlock” that allowed so much progress this past year? I ask as a noob in this area.
- mlboss 3y agoStable diffusion open source release and llama release
- alberth 3y agoBut what technically allowed for so much progress? There’s been open source AI/ML for 20+ years. Nothing comes close to the massive milestones over the past year.
- Chabsff 3y agoPublic availability of large transformer-based foundation models trained at great expense, which is what OP is referring to, is definitely unprecedented.
- kmeisthax 3y agoAttention, transformers, diffusion. Prior image synthesis techniques - i.e. GANs - had problems that made it difficult to scale them up, whereas the current techniques seem to have no limit other than the amount of RAM in your GPU.
- jasonjmcghee 3y agoPeople figuring out how to train and scale newer architectures (like transfomers) effectively, to be wildly larger than ever before. Take AlexNet - the major "oh shit" moment in image classification. It had an absolutely mind-blowing number of parameters at a whopping 62 million. Holy shit, what a large network, right? Absolutely unprecedented. Now, for language models, anything under 1B parameters is a toy that barely works. Stable diffusion has around 1B or so - or the early models did, I'm sure they're larger now. A whole lot of smart people had to do a bunch of cool stuff to be able to keep networks working at all at that size. Many, many times over the years, people have tried to make larger networks, which fail to converge (read: learn to do something useful) in all sorts of crazy ways. At this size, it's also expensive to train these things from scratch, and takes a shit-ton of data, so research/discovery of new things is slow and difficult. But, we kind of climbed over a cliff, and now things are absolutely taking off in all the fields around this kind of stuff. Take a look at XTTSv2 for example, a leading open source text-to-speech model. It uses multiple models in its architecture, but one of them is GPT. There are a few key models that are still being used in a bunch of different modalities like CLIP, U-Net, GPT, etc. or similar variants. When they were released / made available, people jumped on them and started experimenting.
- hanniabu 3y ago> but the real utility of this will be the temporal consistency The main utility will me misinformation
- deleted 3y ago[deleted]
- kornesh 3y agoYeah, solving the flickering problem and achieving temporal consistency will be the key to realize the full potential of generative video models. Right now, AnimateDiff is leading the way in consistency but I'm really excited to see what people will do with this new model.