16 ms·
Stable Video Diffusion
- deleted 3y ago[deleted]
- minimaxir 3y agoModel weights (two variations, each 10GB) are available without waitlist/approval: https://huggingface.co/stabilityai/stable-video-diffusion-img2vid-xt https://huggingface.co/stabilityai/stable-video-diffusion-im... The LICENSE is a special non-commercial one: https://huggingface.co/stabilityai/stable-video-diffusion-img2vid-xt/blob/main/LICENSE https://huggingface.co/stabilityai/stable-video-diffusion-im... It's unclear how exactly to run it easily: diffusers has video generation support now but need to see if it plugs in seamlessly.
- ronsor 3y agoRegular reminder that it is very likely that model weights can't be copyrighted (and thus can't be licensed).
- deleted 3y ago[deleted]
- chankstein38 3y agoIt looks like the huggingface page links their github that seems to have python scripts to run these: https://github.com/Stability-AI/generative-models https://github.com/Stability-AI/generative-models
- minimaxir 3y agoThose scripts aren't as easy to use or iterate upon since they are CLI apps instead of a REPL like a Colab/Jupyter Notebook (although these models probably will not run in a normal Colab without shenanigans). They can be hacked into a Jupyter Notebook but it's really not fun.
- valine 3y agoThe rate of progress in ML this past year has been breath taking. I can’t wait to see what people do with this once controlnet is properly adapted to video. Generating videos from scratch is cool, but the real utility of this will be the temporal consistency. Getting stable video out of stable diffusion typically involves lots of manual post processing to remove flicker.
- Der_Einzige 3y agoControlnet is adapted to video today, the issues are that it's very slow. Haven't you seen the insane quality of videos on civitai?
- valine 3y agoI have seen them, the workflows to create those videos are extremely labor intensive. Control net lets you maintain poses between frames, it doesn’t solve the temporal consistency of small details.
- mattnewton 3y agoPeople use animatediff’s motion module (or other models that have cross frame attention layers). Consistency is close to being solved.
- valine 3y agoHopefully this new model will be a step beyond what you can do with animatediff
- dragonwriter 3y agoTemporal consistency is improving, but “close to being solved” is very optimistic.
- mattnewton 3y agoNo I think we’re actually close. My source is I’m working on this problem and the incredible progress of our tiny 3 person team at drip.art (http://api.drip.art http://api.drip.art) - we can generate a lot of frames that are consistent, and with interpolation between them, smoothly restyle even long videos. Cross-frame attention works for most cases, it just needs to be scaled up. And that’s just for diffusion focused approaches like ours. There are probably other techniques from the token flow or nerf family of approaches close to breakout levels of quality, tons of talented researchers working on that too.
- ericpauley 3y agoI'm still puzzled as to how these "non-commercial" model licenses are supposed to be enforceable. Software licenses govern the redistribution of the software, not products produced with it. An image isn't GPL'd because it was produced with GIMP.
- cubefox 3y agoNobody claimed otherwise?
- littlethoughts 3y agoFantasy.ai was subject to controversy for attempting to license models.
- not2b 3y agoThere are sites that make Stable Diffusion-derived models available, along with GPU resources, and they sell the service of generating images from the models. The company isn't permitting that use, and it seems that they could find violators and shut them down.
- Der_Einzige 3y agoThey're not enforceable.
- yorwba 3y agoThe license is a contract that allows you to use the software provided you fulfill some conditions. If you do not fulfill the conditions, you have no right to a copy of the software and can be sued. This enforcement mechanism is the same whether the conditions are that you include source code with copies you redistribute, or that you may only use it for evil, or that you must pay a monthly fee. Of course this enforcement mechanism may turn out to be ineffective if it's hard to discover that you're violating the conditions.
- comex 3y agoIt also somewhat depends on open legal questions like whether models are copyrightable and, if so, whether model outputs are derivative works of the model. Suppose that models are not copyrightable, due to their not being the product of human creativity (this is debatable). Then the creator can still require people to agree to contractual terms before downloading the model from them, presumably including the usage limitations as well as an agreement not to redistribute the model to anyone else who does not also agree. Agreement can happen explicitly by pressing a button, or potentially implicitly just by downloading the model from them, if the terms are clearly disclosed beforehand. But if someone decides on their own (not induced by you in any way) to violate the contract by uploading it somewhere else, and you passively download it from there, then you may be in the clear.
- helpmenotok 3y agoCan this be used for porn?
- theodric 3y agoIf it can't, someone will massage it until it can. Porn, and probably also stock video to sell to YouTubers.
- _qxjp 3y agoVery unusual comment. I do not think so as the chance of constructing a fleshy eldritch horror is quite high.
- tstrimple 3y ago> I do not think so as the chance of constructing a fleshy eldritch horror is quite high. There is a market for everything!
- johndevor 3y agoHow is that not the first question to ask? Porn has proven to be a fantastic litmus test of fast market penetration when it comes to new technologies.
- _qxjp 3y agoThis is true. I was hoping my educated guess of the outcome would minimize the possibility of anyone attempting this. And yet, here we are - the only losing strategy in the technology sector is to not try at all.
- throwaway743 3y agoNo pun intended?
- xanderlewis 3y agoMarket what?
- crtasm 3y ago
- christkv 3y agoLooks like I'm still good for my bet with some friends that before 2028 a team of 5-10 people will create a blockbuster style movie that today costs 100+ million USD on a shoestring budget and we won't be able to tell.
- CamperBob2 3y agoIt'll happen, but I think you're early. 2038 for sure, unless something drastic happens to stop it (or is forced to happen.)
- accrual 3y agoThe first full-length AI generated movie will be an important milestone for sure, and will probably become a "required watch" for future AI history classes. I wonder what the Rotten Tomatoes page will look like.
- throwaway743 3y agoDefinitely a big first for benchmarks. After that hyper personalized content/media generated on-demand
- ben_w 3y agoI wouldn't bet either way. Back in the mid 90s to 2010 or so, graphical improvements were hailed as photorealistic only to be improved upon with each subsequent blockbuster game. I think we're in a similar phase with AI[0]: every new release in $category is better, gets hailed as super fantastic world changing, is improved upon in the subsequent Two Minute Papers video on $category, and the cycle repeats. [0] all of them: LLMs, image generators, cars, robots, voice recognition and synthesis, scientific research, …
- btbuildem 3y agoIn the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue jays" would be near "toronto" or "cn tower". The improvements in scale and speed (image -> now video) are impressive, but given how incredibly able the image generation models are, they simultaneously feel crippled and limited by their lack of editing / iteration ability. Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close.
- appplication 3y agoI don’t spend a lot of time keeping up with the space, but I could have sworn I’ve seen a demo that allowed you to iterate in the way you’re suggesting. Maybe someone else can link it.
- accrual 3y agoIt's not exactly like GP described (e.g. move bike to the left) but there is a more advanced SD technique called inpainting [0] that allows you to manually recompose parts of the image, e.g. to fix bad eyes and hands. [0] https://stable-diffusion-art.com/inpainting_basics/ https://stable-diffusion-art.com/inpainting_basics/
- ssalka 3y agoMy guess is you're thinking of InstructPix2Pix[1], with prompts like "make the sky green" or "replace the fruits with cake" [1] https://github.com/timothybrooks/instruct-pix2pix https://github.com/timothybrooks/instruct-pix2pix
- appplication 3y agoThis is exactly it!
- dinvlad 3y agoSeems relatively unimpressive tbh - it's not really a video, and we've seen this kind of thing for a few months now
- accrual 3y agoIt seems like the breakthrough is that the video generating method is now baked into the model and generator. I've seen several fairly impressive AI animations as well, but until now, I assumed they were tediously cobbled together by hacking on the still-image SD models.
- youssefabdelm 3y agoCan't wait for these things to not suck
- accrual 3y agoIt's definitely pretty impressive already. If there could be some kind of "final pass" to remove the slightly glitchy generative artifacts, these look completely passible for simple .gif/.webm header images. Especially if they could be made to loop smoothly ala Snapchat's bounce filter.
- accrual 3y agoFascinating leap forward. It makes me think of the difference between ancestral and non-ancestral samplers, e.g. Euler vs Euler Ancestral. With Euler, the output is somewhat deterministic and doesn't vary with increasing sampling steps, but with Ancestral, noise is added to each step which creates more variety but is more random/stochastic. I assume to create video, the sampler needs to lean heavily on the previous frame while injecting some kind of sub-prompt, like rotate <object> to the left by 5 degrees, etc. I like the phrase another commenter used, "temporal consistency". Edit: Indeed the special sauce is "temporal layers". [0] > Recently, latent diffusion models trained for 2D image synthesis have been turned into generative video models by inserting temporal layers and finetuning them on small, high-quality video datasets [0] https://stability.ai/research/stable-video-diffusion-scaling-latent-video-diffusion-models-to-large-datasets https://stability.ai/research/stable-video-diffusion-scaling...
- adventured 3y agoThe hardest problem the Stable Diffusion community has dealt with in terms of quality has been in the video space, largely in relation to the consistency between frames. It's probably the most commonly discussed problem for example on r/stablediffusion. Temporal consistency is the popular term for that. So this example was posted an hour ago, and it's jumping all over the place frame to frame (somewhat weak temporal consistency). The author appears to have used pretty straight-forward text2img + Animatediff: https://www.reddit.com/r/StableDiffusion/comments/180no09/on_the_swing_animatediff_text2image_only/ https://www.reddit.com/r/StableDiffusion/comments/180no09/on... Fixing that frame to frame jitter related to animation is probably the most in-demand thing around Stable Diffusion right now. Animatediff motion painting made a splash the other day: https://www.reddit.com/r/StableDiffusion/comments/17xnqn7/roll_your_own_motion_brush_with_animatediff_and/ https://www.reddit.com/r/StableDiffusion/comments/17xnqn7/ro... It's definitely an exciting time around SD + animation. You can see how close it is to reaching the next level of generation.
- torginus 3y agoI admit I'm ignorant about these model's inner workings, but I don't understand why text is the chosen input format for these models. It was the same for image generation, where one needed to produce text prompts to create the image, and stuff like img2img and Controlnet that allowed things like controlling poses and inpainting, or having multiple prompts with masks controlling which part of the image is influenced by which prompt.
- gorbypark 3y agoAccording to the GitHub repo this is an "image-to-video model". They tease of an upcoming "text to video" interface on the linked landing page, though. My guess is that interface will use a text-to-image model and then feed that into the image-to-video model.
- pizzafeelsright 3y agoImago Deo? The Word is what is spoken when we create. The input eventually becomes meanings mapped to reality.
- awongh 3y agoIt makes sense that they had to take out all of the cuts and fades from the training data to improve results. I’m the background section of the research paper they mention “temporal convolution layers”, can anyone explain what that is? What sort of training data is the input to represent temporal states between images that make up a video? Or does that mean something else?
- machinekob 3y agoI would assume is something similar to joining multiple frames/attentions? in channel dimension and then moving values inside so convolution will have access to some channels from other video frames. I was working on similar idea few years ago using this paper as reference and it was working extremely well for consistency also helping with flicker. https://arxiv.org/abs/1811.08383 https://arxiv.org/abs/1811.08383
- karelpeeters 3y agoIt means that instead of (only) doing convolution in spatial dimensions, it also(/instead) happens in the temporal dimension. A good resource for the "instead" case: https://unit8.com/resources/temporal-convolutional-networks-and-forecasting/ https://unit8.com/resources/temporal-convolutional-networks-... The "also" case is an example of 3D convolution, an example of a paper that uses it: https://www.cv-foundation.org/openaccess/content_iccv_2015/papers/Tran_Learning_Spatiotemporal_Features_ICCV_2015_paper.pdf https://www.cv-foundation.org/openaccess/content_iccv_2015/p...
- spaceman_2020 3y agoA seemingly off topic question, but with enough compute and optimization, could you eventually simulate “reality”? Like, at this point, what are the technical counters to the assertion that our world is a simulation?
- refulgentis 3y agoA little too freshman's first bit off a bong for me. There is, of course, substantial differences between video and reality. Let's steel-man — you mean 3D VR. Let's stipulate there's a headset today that renders 3D visually indistinguishable from reality. We're still short the other 4 senses Much like faith, there's always a way to sort of escape the traps here and say "can you PROVE this is base reality" The general technical argument against "brain in a vat being stimulated" would be the computation expense of doing such, but you can also write that off with the equivalent of foveated rendering but for all senses / entities
- 2-718-281-828 3y ago> Like, at this point, what are the technical counters to the assertion that our world is a simulation? How about this theory is neither verifiable nor falsifiable.
- vidarh 3y agoThe general concept is not falsifiable, but many variations might be, or their inverse might be. E.g. the theory that we are not in a simulation would in general be falsifiable by finding an "escape" from a simulation and so showing we are in one (but not finding an escape of course tells us nothing). It's not a very useful endeavour to worry about, but it can be fun to speculate about what might give rise to testable hypotheses and what that might tell us about the world.
- deleted 3y ago[deleted]
- tracerbulletx 3y ago
- epiccoleman 3y agoThis is really, really cool. A few months ago I was playing with some of the "video" generation models on Replicate, and I got some really neat results[1], but it was very clear that the resulting videos were made from prompting each "frame" with the previous one. This looks like it can actually figure out how to make something that has a higher level context to it. It's crazy to see this level of progress in just a bit over half a year. [1]: https://epiccoleman.com/posts/2023-03-05-deforum-stable-diffusion https://epiccoleman.com/posts/2023-03-05-deforum-stable-diff...
- richthekid 3y agoThis is gonna change everything
- jetsetk 3y agoIs it? How so?
- Chabsff 3y agoIt's really not. Don't get me wrong, this is insanely cool, but it's still a long way from good enough to be truly disruptive.
- echelon 3y agoOne year. All of Hollywood falls.
- Chabsff 3y agoNo offense, but this is absolutely delusional. As long as people can "clock" content generated from these models, it will be treated by consumers as low-effort drivel, no matter how much actual artistic effort goes in the exercise. Only once these systems push through the threshold of being indistinguishable from artistry will all hell break loose, and we are still very far from that. Paint-by-numbers low-effort market-driven stuff will take a hit for sure, but that's only a portion of the market, and frankly not one I'm going to be missing.
- ben_w 3y agoVery far, yes, but also in a fast moving field. CGI in films used to be obvious all the time no matter how good the artists using it, now it's everywhere and only noticeable when that's the point; the gap from Tron to Fellowship of the Ring was 19.5 years. My guess is the analogy here puts the quality of existing genAI somewhere near the equivalent of early TV CGI, given its use in one of the Marvel title sequences etc., but it is just an analogy and there's no guarantees of anything either way.
- nbzso 3y agoModel chain: Instance One : Act as a top tier Hollywood scenarist, use the public available data for emotional sentiment to generate a storyline, apply the well known archetypes from proven blockbusters for character development. Move to instance two. Instance Two: Act as top tier producer. {insert generated prompt}. Move to instance three. Instance Three: Generate Meta-humans and load personality traits. Move to instance four. Instance Four: Act as a top tier director.{insert generated prompt}. Move to instance five. Instance Five: Act as a top tier editor.{insert generated prompt}. Move to instance six. Instance Six: Act as a top tier marketing and advertisement agency.{insert generated prompt}. Move to instance seven. Instance Seven: Act as a top tier accountant, generate an interface to real-time ROI data and give me the results on an optimized timeline into my AI induced dream. Personal GPT: Buy some stocks, diversify my portfolio, stock up on synthetic meat, bug-coke and Soma. Call my mom and tell her I made it.
- aliljet 3y agoI've been following this space very very closely and the killer feature would be to be able to generate these full featured videos for longer than a few seconds with consistently shaped "characters" (e.g., flowers, and grass, and houses, and cars, actors, etc.). Right now, it's not clear to me that this is achieving that objective. This feels like it could be great to create short GIFs, but at what cost? To be clear, this remains wicked, wicked, wicked exciting.
- speedgoose 3y agoHas anyone managed to run the thing? I got the streamlit demo to start after fighting with pytorch, mamba, and pip for half an hour, but the demo runs out of GPU memory after a little while. I have 24GB on GPU on the machine I used, does it need more?
- mkaic 3y agoHave heard from others attempting it that it needs 40GB, so basically an A100/A6000/H100 or other large card. Or an Apple Silicon Mac with a bunch of unified memory, I guess.
- mlboss 3y agoGive it a week.
- speedgoose 3y agoAlright thanks for the information. I will try to justify using one A100 for my "very important" research activities.
- skonteam 3y agoYeah, got a 24GB 4090, try to reduce the number of frames decoded to something like 4 or 8. Although, keep in mind it caps the 24Gb and goes to RAM (with the latest nvidia drivers).
- speedgoose 3y agoOh yes it works, thanks!
- nwoli 3y agoIs the checkpoint default fp16 or fp32?
- neaumusic 3y agoIt's funny that still don't really have video wallpapers on most devices (I'm only aware of Wallpaper Engine on Windows)
- Sohcahtoa82 3y agoI had a video wallpaper on my Motorola Droid back in 2010.
- tetris11 3y agoand a battery life of...? I do wonder if there have been any codec studies that measure power usage with respect to RAM
- spupy 3y agoMplayer/MPV used to be able to play videos in the X root window like a wallpaper. No idea if it still works nowadays.
- pcj-github 3y agoSoon the hollywood strike won't even matter, won't need any of those jobs. Entire west coast economy obliterated.
- jonplackett 3y agoIs this available in the stability API any time soon?
- chrononaut 3y agoMuch like in static images, the subtle unintended imperfections are quite interesting to observe. For example, the man in the cowboy hat seems he is almost gagging. In the train video the tracks seem to be too wide while the train ice skates across them.
- shaileshm 3y agoThis field moves so fast. Blink an eye and there is another new paper. This is really cool and the learning speed of us humans is insane! Really excited on using it for downstream tasks! I wonder how easy it is to integrate animatediff with this model? Also, can someone benchmark it on m3 devices? It would be cool to see if it is worth getting on to run these diffusion inferences and development. If m3 pro can allow finetuning it would be amazing to use it on downstream tasks!
- nuclearsugar 3y agoVery excited to play with this. Some of my latest experiments - https://www.jasonfletcher.info/vjloops/ https://www.jasonfletcher.info/vjloops/
- rbhuta 3y agoWe're hosting this free (no credit card needed) at https://app.decoherence.co/stablevideo https://app.decoherence.co/stablevideo Disclaimer: Google log-in required to help us reduce spam. Let me know what you think of it! It works best on landscape images from my tests.
- didip 3y agoStability.ai, please make sure your board is sane.
- RandomBK 3y agoNeeds 40GB VRAM, down to 24GB by reducing the number of frames processed in parallel.
- firefoxd 3y agoI understand the magnitude of innovation that's going on here. But still feel like we are generating these videos with both hands tied behind our backs. In other words, it's nearly impossible to edit the videos in this constraints. (Imagine trying to edit the blue Jays to get the perfect view). Since videos are rarely consumed raw, what if this becomes a pipeline in Blender instead? (Blender the 3d software). Now the video becomes a complete scene with all the key elements of the text input animated. You have your textures, you have your animation, you have your camera, you have all the objects in place. We can even have the render engine in the pipeline to increase the speed of video generation. It may sound like I'm complaining, but I'm just ask making a feature request...
- huytersd 3y agoWhat would solve all these issues is full generation of 3D models that we hopefully get a chance to see over the next decade. I’ve been advocating for a solid LiDAR camera on the iPhone so there is a lot of training data for these LLMs.
- ricardobeat 3y ago> I’ve been advocating for a solid LiDAR camera on the iPhone What do you mean by “advocating”? The iPhone has had a LiDAR camera since 2020.
- jwoodbridge 3y agowe're working on this - dream3d.com
- Eduard 3y agocannot join the waiting list (nor opt in for marketing newsletter), because the sign-up form checkboxes don't toggle on android mobile Chrome or Firefox.
- gregorymichael 3y agoHow long until Replicate has this available?
- rbhuta 3y agoWe're hosting this free (no credit card needed) at https://app.decoherence.co/stablevideo https://app.decoherence.co/stablevideo Disclaimer: Google log-in required to help us reduce spam. Let me know what you think of it! It works best on landscape images from my tests.
- radicality 3y agoLooks like there is a WIP here: https://replicate.com/lucataco/svd https://replicate.com/lucataco/svd
- rbhuta 3y agoVRAM requirements are big for this launch. We're hosting this for free at https://app.decoherence.co/stablevideo https://app.decoherence.co/stablevideo. Disclaimer: Google log-in required to help us reduce spam.
- xena 3y agoHow big is big?
- whywhywhywhy 3y ago40GB although hearing reports 3090 can do low frame counts
- zvictor 3y agoit's worth paying your subscription just for these free videos. would those have the watermark removed if I go "Basic"?
- keiferski 3y agoQuestion for anyone more familiar with this space: are there any high-quality tools which take an image and make it into a short video? For example, an image of a tree becomes a video of a tree swaying in the wind. I have googled for it but mostly just get low quality web tools.
- devdiary 3y agoA default glitch effect in the video can make the distortions a "feature not a bug"
- renlo 3y agoHow much longer will it be until we can play "video games" which consist of user-input streamed to an AI that generates video output and streams it to the player's screen?
- slow_numbnut 3y agoIf you're willing to accept text based output then Text adventure style games and even simulating bash was possible using chatgpt until openAI nerfed it.
- AltruisticGapHN 3y agoThese are basically like animated postcards, like you often see now on loading screens in videogames. A single picture has been animated. Still a long shot from actual video.
- siddbudd 3y ago"2 more papers down the line"...
- rvion 3y agoFinally ! Now that this is out, I can finally start adding proper video widgets to CushyStudio https://github.com/rvion/CushyStudio#readme https://github.com/rvion/CushyStudio#readme . Really hope I can get in touch with StabilityAi people soon. Maybe Hacker News will help
- TruthWillHurt 3y agoAnd thanks to the porn community on Civit.ai!
- iamgopal 3y agoVery soon, we will be able to change story line of a web series dynamically, a little more thrill, a little more comedy, changing character face to matching ours and others, all in 3D with 360 degree view, how far are we from this ? 5 year ?
- niek_pas 3y agoAt least several decades, I’d say. This is a hugely complex, multifaceted problem. LLMs can’t even write half-decent screenplays yet.
- LoveMortuus 3y agoOnce text-to-video is good enough and once text generation is good enough, we could legit actually have endless TV shows produced by individuals! We're probably still far away from that, but it is exciting to think about! I think this will really open new ways and new doors to creativity and creative expression.