9 ms·
Stable Diffusion animation
- Dwedit 4y agoIn this example, 25 frames are generated using Stable Diffusion, then frames are interpolated using FILM-Net. I hadn't see FILM-net before, it looks really neat.
- o_____________o 4y agoThis? Looks pretty amazing: https://film-net.github.io https://film-net.github.io https://github.com/google-research/frame-interpolation https://github.com/google-research/frame-interpolation
- bfirsh 4y agoYou can run it on Replicate too! https://replicate.com/google-research/frame-interpolation https://replicate.com/google-research/frame-interpolation
- eezurr 4y agoNeat. Didnt Microsoft release a tool that morphed between two photos a decade ago? I dont recall the name unfortunately, but the effect was similar, less quality.
- nopenopenopeno 4y agoI was doing this with an After Effects plugin called Twixtor long before a decade ago. https://revisionfx.com/products/twixtor/ https://revisionfx.com/products/twixtor/
- Jaruzel 4y agoBack in the DOS days I had (and a still do) a copy of a tool called 'Morph', it took two GIF images, and you placed marker points on the first image and then again on the second, and it generated an MPEG2 file of the morphing. Very impressive for the time.
- dr_dshiv 4y agoMid 90s, baby!
- Dwedit 4y agoI had the "DMORPH" program, which had you construct a grid, and it would generate separate image for every frame. Then you had to use "DTA" to turn it into a FLIC file. No MPEG here, no animated .gif here, you had FLIC as your animation format.
- Jaruzel 4y agoFLIC animation, now that is a blast from the past!
- maven29 4y agoYou're thinking of Microsoft PhotoSynth from 2006. This is similar to what google streetview showed up with. From what I can see, they later revamped PhotoSynth to include actual 3d mesh reconstruction in 2014.
- nl 4y agoMorphing is different to interpolation and has been available on consumer computers since the 1990s (I remember making gifs in 1996 with it). In morphing you end up with a blury mess for the between frames. This technique tracks individual features (eg eyes) and keeps them coherent.
- mrpf1ster 4y agoI feel like I’m watching an explosion of progress in AI image generation in real-time. Every day there’s a new application of Stable Diffusion. It’s incredible to watch unfold
- zmmmmm 4y agoI think it's pretty notable how most of the explosion happened since stable diffusion released their model and code as open source, while Dall-E generated initial excitement with their closed source model, but limited progress / creativity since. It's a pretty nice demonstration I think of how much innovation can happen from openness.
- kristopolous 4y agoYes. Example #5*10^7 or so. Some people are just opposed ideologically or due to temperament. Locking things down is one of the best ways to make them die
- napolux 4y agoGPT-4 (or a reduced version of it) should be opensource too if you ask me :P
- deleted 4y ago[deleted]
- ducktective 4y agoIncredible for you, depressing for me. I feel left behind...
- deleted 4y ago[deleted]
- r3trohack3r 4y agoI just came across this on twitter, every frame appears to be an evolution of the previous frame using img2img paired with a tilt/zoom to create a psychedelic animation. The author claims to have made this with Stable Diffusion, Disco, and Wiggle: https://www.youtube.com/watch?v=Nz_n0qxqoPg https://www.youtube.com/watch?v=Nz_n0qxqoPg I believe Wiggle is used to automate the tilt/zoom between frames.
- westoncb 4y agoInteresting to see in comparison to those old latent space zoom vids making the rounds in like ~2014..
- alx__ 4y agoThat is wild, feels like an intense dream you can't wake up from
- actionfromafar 4y agoI like their earlier work more, with the audio pictograms.
- deleted 4y ago[deleted]
- waking_at_night 4y agoThats really great and has come a long way since the beginnings... I had this video animation going on back in the day when VQGAN was still all the rave: https://www.youtube.com/watch?v=CgDbbg802-8 https://www.youtube.com/watch?v=CgDbbg802-8 Its incredible what 6 months only can do
- alhirzel 4y agoReminds me of Electric Sheep[1] - I can't wait until someone hooks up news stories to have abstract news as a screensaver! Also reminds me of a time lapse of someone painting - like [2] [1]: https://electricsheep.org/ https://electricsheep.org/ [2]: https://www.youtube.com/watch?v=2O4ccHgfcl8 https://www.youtube.com/watch?v=2O4ccHgfcl8
- fragmede 4y agoThere's also https://gist.github.com/karpathy/00103b0037c5aaea32fe1da1af553355 https://gist.github.com/karpathy/00103b0037c5aaea32fe1da1af5...
- autoexec 4y agoAI in animation has been interesting to me for a while now. It leaves me a little conflicted though. If we get to the point where we can throw key drawings at AI and let it handle all the inbewteens without a bunch of tweaking and cleanup afterwards it's going to really suck for places like Korea! I guess all those inbeatweeners will just be another victim of automation. I've always loved animation, but I'll admit part of that comes from the hubris involved. It's pure insanity that people ever drew, by hand, mountains of individual drawings each slightly changed and assembled them into compelling illusions to tell stories. The amount of work that goes into animation is just staggering and anyone sensible would have rejected the entire concept as absurd. I wonder if animation will start losing part of its magic for me when it's done primarily by AI. On the other hand though, another thing I've always loved about animation as a storytelling medium is that it isn't as limited by practical concerns like physics or reality. If something can be imagined, it can be drawn and animated if somebody has the skill and the resources to fund the massive amounts of work. It's time/money that forces animators to take shortcuts and make compromises. Creative decisions are made and rejected all the time due to those constraints. If AI driven animation gets more advanced to the point where that's no longer such a barrier it could create output more in line with the vision of creators and that's exciting too! I hope that traditional hand drawn animation never dies, but I look forward to seeing how AI continues to change the industry and the output.
- megablast 4y ago> leaves me a little conflicted though. If we get to the point where we can throw key drawings at AI and let it handle all the inbewteens without a bunch of tweaking and cleanup afterwards it's going to really suck for places like Korea! These comments on every single post are getting really boring.
- vasco 4y agoYou can probably expect them for any interesting technology forever into the future, since people have made these useless complaints for hundreds of years at least.
- Lwepz 4y agoIt's clear that the next frontier is to have 3D-space instead of image space transitions. Language itself is very static and action verbs are not enough to specify scene dynamics. I suppose we would need: A. an enriched version of natural language that refines the dynamic processes that occur in a scene B. a data set of isolated processes labeled in the language described in A. I've had a hard time finding ongoing work on A. and B, perhaps it isn't much of a priority for research groups.
- redox99 4y agoFor 3D we would probably need something like Blender or similar, because at some point it's just easier to use a 3D software to pinpoint where you want stuff to be, than try to use words. Imagine opening blender and typing > A medium sized classroom, well lit, with two blackboards and many geography posters And the AI just generates all the 3D meshes and places them appropriately. Repeat that for other props or characters that you need. After that you can manually tweak the scene as you currently would (moving things, etc). Then you select a character, and to animate you tell it > The character calmly walks to the door and proceeds to open it You could literally do a 100+ hour job in 5 minutes.
- danwills 4y agoI think you're right about the possibilities here and I love the idea and have thought similar things myself too, but to me the inclusion of a 3D element should probably be a format, not necessarily locked into any specific app such as Blender. Maybe (Pixar) USD is one possible format that could be used for general 3D interchange for this kind of thing?
- redox99 4y agoThe blender thing was just as an example. Of course it would be possible to convert whatever the output of the model is, to whatever software you want to plug it in.
- NoMoreBro 4y agoVery cool! I generated this with 1000 images https://twitter.com/UnshushProject/status/1563158214577094657 https://twitter.com/UnshushProject/status/156315821457709465... using the Deforum's Colab[1], it's really easy and now has interpolation too. It was the very first video, I could have made something great but, you know, awesome guys keep releasing AI tech and I'm like a child at Luna Park right now, not able to concentrate. If you are interested in my project (I doubt, you are too busy playing like me) I'm posting a lot of things on https://unshush.com https://unshush.com and on the Instagram account: https://www.instagram.com/unshushproject/ https://www.instagram.com/unshushproject/ (Sorry for posting my stuff but I'm not very social so no one will ever see them) If you want to generate videos I can share some links I bookmarked of software/code to make them more smooth. [1] Deforum's Colab (based on Stable Diffusion): https://colab.research.google.com/github/deforum/stable-diffusion/blob/main/Deforum_Stable_Diffusion.ipynb https://colab.research.google.com/github/deforum/stable-diff...
- yreg 4y agoI'm definitely interested in your video related bookmarks. You have generated some pretty cool designs.
- NoMoreBro 4y agoSure! This would be my approach (and tools) if I was smarter: If you make the generations with some similarities and use the right interpolation, you don't need 1000 images like my video and can obtain a smooth movement. First, generate images with some kind of visual anchor (background, an object). You can use frames generated using the previous frame as reference image, or the same seed but different prompt/parameters, or you can go wild using img2img/inpainting (btw I struggle to find an inpainting tool for Stable Diffusion: they seem to be just img2img with a mask, without contest). Then pass the generated images to one of the most recent interpolation algorithms, like this one https://github.com/megvii-research/ECCV2022-RIFE https://github.com/megvii-research/ECCV2022-RIFE or the one used in the replicate we are commenting on (someone posted this reference: https://github.com/google-research/frame-interpolation https://github.com/google-research/frame-interpolation ) The first link reports some free and paid implementation and a Colab, so depending on how deep you want to go, you have a lot of choices. In the end, I'd use some good app to stabilize the image if needed, to get a more "calm" look. I use Luma Fusion, but it's a paid app (cheap, one-time payment, for iOS). I'm sure there are a ton of open-source implementations. It's an approach similar to the animation on replicate, but it allows a lot of fine-tuning and you can add new animation ideas/tools to the process. Nothing revolutionary, but I hope it helps! > You have generated some pretty cool designs. Thanks! I put in a lot of work in the last weeks. The project has a mission, I wrote something, but it's not ready yet. I believe it will be with the launch of Dall-E 8 :-/
- simultsop 4y agoBrace yourselves, Cartoon AI Network, coming soon xD
- justinlloyd 4y agoLast year, when 3090 GPUs were astronomically priced, I thought "screw it, I'll just buy an RTX A5000 for a couple of hundred bucks more." Which begat a second A5000 for "reasons." It was almost prescient. Now all these models are coming out requiring slightly higher VRAM GPUs than a 3090, i.e. more in the range of the A5000, and I get to run them. I am a kid in a candy store this past couple of weeks.
- rjh29 4y agoThe A5000 and 3090 both have 24GB of ram?
- dannyw 4y agoMaybe GP meant 3080.
- rjh29 4y agoFor anyone interested in a GPU btw, the 3090 TI had a huge price cut and costs only a bit more than the 3090 right now.
- CuriouslyC 4y agoYup! I've never splurged on a GPU before, but a 3090 TI lets you do textual inversion, and fine tune GPT-J (neither of which you can do with <24gb vram), so now I've got one sitting on my study floor waiting to be installed :)
- tough 4y ago> https://twitter.com/dreamwieber/status/1565008078466326528?s https://twitter.com/dreamwieber/status/1565008078466326528?s... same mine arrives tomorrow, can't wait to try textual-inversion
- justinlloyd 4y agoI should have mentioned linking the GPUs. I meant requiring slightly more VRAM than a single 3090 can handle. Or a 3080 as others have pointed out. The main difference between the way the 3090 works and the A5000 works is SLI vs NvLink/NvSwitch. I believe the 3090 uses NvLink, but not quite in the same way the A5000 does. I can chain together far more A5000's than I can 3090's. Eight A500's vs four 3090's IIRC. And even chain them across machines with the right h/w, though that's probably a bit of a stretch for my budget. Also, the A5000 will share VRAM, giving me a total usable heap of 48GB with two cards, whereas the 3090 will be limited to 24GB each. I can also share the A5000's with multiple VMs simultaneously, whereas with the 3090's I am stuck doing GPU pass through. All that for only a couple of percentage points drop in performance in video games.
- t2hv33 4y agoOk
- fagerhult 4y agoAndreas, author of the Replicate model here -- though "author" feels wrong since I basically just stitched two amazing models together. The thing that really strikes me is that open source ML is starting to behave like open source software. I was able to take a pretrained text-to-image model and combine it with a pretrained video frame interpolation model and the two actually fit together! I didn't have to re-train or fine tune or map between incompatible embedding spaces, because these models can generalize to basically any image. I could treat these models as modular building blocks. It just makes your creative mind spin. What if you generate some speech with https://replicate.com/afiaka87/tortoise-tts https://replicate.com/afiaka87/tortoise-tts, generate an image of an alien with Stable Diffusion, and then feed those two into https://replicate.com/wyhsirius/lia https://replicate.com/wyhsirius/lia. Talking alien! Machine learning is starting to become really fun, even if you don't know anything about partial derivatives.
- NaturalPhallacy 4y ago>Talking alien! Maybe that's what we were supposed to do all along. Not find or be found by aliens, but to invent them.
- rnjesus 4y agothat’s exactly what we’ve been doing
- noduerme 4y agoIt's a nifty piece of work. Often when you're trying to get an answer from a regression model or a neural net you have to try to craft your inputs so carefully that you already sort of know, intuitively, what it will figure out. In some way the thought process of the refining the input is more valuable in a lot of quantitative cases than the actual output. This is simply very impressive... whether or not it was humbly stitched together, you were sort of the first to do it, so take pride. The next real magic will be reading its net and figuring out how to get [vfx/film] effects from it... which if I were you would probably occupy 22 hours of my day now.
- benreesman 4y ago
- amelius 4y agoI guess it is missing the physics simulation between frames. Perhaps that is the next big step for ML to get right.
- zone411 4y agoI played a bit with smoothly interpolating between Stable Diffusion prompts and the effect can be pretty cool but it's hard to avoid discontinuities (like the object changing its orientation), even when using some additional tricks like reusing the previous frame as the initial image or generating several new frames and choosing the one that's closest to the previous frame. You basically have to get lucky with the seed. It probably makes most sense to just wait for video models that take temporal consistency into account explicitly or generate 3D models. There is a lot of promising research out there already, so it's just a matter of time.
- deleted 4y ago[deleted]
- desindol 4y agoMaybe I can shine some light on the debate from an concept artist standpoint that works in VFX and advertising. I worked on feature films (3 of them in the imdb top 100), tv shows (like game of thrones) and hundreds of AD campaigns. In the last 10 year the work of a concept artist changed dramatically we have gone from purely painted concept art to mostly "photobashed". Photobashed means basically that you rip apart other images and stitch them together to get the desired image. Some start with a rough sketch for the composition or make really rough grey shade 3d model and "overpaint" them. When it comes to "photobashing" the disregard for copyrights was always there and it's the worst in smaller studios and a bit better in the leading ones. Still most of the time everyone argues that if you only use really small parts of the images it is covered by fair use. There are some examples were studios got sued but mostly without bigger financial impact. A few months ago I started working with "DiscoDiffusion" to generate the images I use to photobash. "DiscoDiffusion" can produce great "painterly" images but struggles with photorealism and is slower, not as coherent as "StableDiffusion". Still the adoption rate in the concept art community was insanely fast. This all got topped by "StableDiffusion" in the last week. Ofc there are still people that want to do it the "right" way and not use AI but we had the same discussion years ago when "photobashing" came into place and some artists still wanted to paint the whole image. As concept artist you are mostly paid for your design thinking that means it is less about the process and more about the finished product. The turnaround time for styleframes got reduced from 3-4 hours while painting to 45 min - 1 hour when photobashing with stable diffusion me and my peers in the studio are now at 20-45 min per styleframe. When "photobashing" most people constrain themselves on their image library and ressources like Photobashing Kits. Not only does "StableDiffusion" cut the time in half it also gives greater freedom in composition and design especially if you are using img2img. So where does this leave us? For the work in fast paced art environments like VFX, games, conept art or advertising "StableDiffusion" is a welcome gamechanger. Tradionalists and Artists outside of the industry might feel threatened but for us in these industries it's a god send.
- shlip 4y agoI think Tradionalists and Artists do what they do because they enjoy the process of creating/creating something unique, not because they want to be productive/fast. So yeah, StableDif is great for commercially constrained environments, plus, it frees some of your time so you can enjoy creating art that matters to you :p
- imhoguy 4y agoAmazing! So counting years now when a first AI feature film hits the cinema screens. Source code: screenplay text. I imagine it may look a bit like "A Scanner Darkly" (2006).
- shlip 4y agoGreat, Miyazaki and Lasseter can finally retire.
- cerol 4y agoMy predictions for 10 - 15 years: - Mandela effect for famous art pieces: "Monalisa was AI generated" "No it wasn't" "Yes it was". - Art critics will get the last laugh, as people start giving them truckloads of money to ask whether a piece of art is human or AI generated.
- ducktective 4y ago>Art critics will get the last laugh, as people start giving them truckloads of money to ask whether a piece of art is human or AI generated. - Each image or painting should have a reference to the organization maintaining it - The organization produces hashes for its artworks - Users compare hashes - Problem solved
- monkeydust 4y agoIt is incredible how fast this thing is progressing. Amazing what you can do with some very smart people and 4,000 A100 cluster ! What is getting very clear though and this link proves it out is that 'prompt engineering' is really a thing, I tried this out and it took a while to get something I would consider half decent. I feel like there is a space here for tools / technologies to 'suggest' prompts based on understanding user intentions. If anyone is actually working on this then reach out to me. Email on profile.
- hackerlight 4y agoWill models like Stable Diffusion be useful for self-driving car research? Like you've got this large NN with weights that are useful for this vision-adjacent task, it should have learned concepts such as edge detection, which could serve as pretrained weights for a self-driving NN?
- stefantalpalaru 4y ago
- synu 4y agoWow, this is incredible. AI tech has been so interesting to follow along lately. Is there something like an index of cool new AI projects that is easy to follow? HN works for this to an extent but I’d love to track more closely.
- avocado2 4y ago
- deleted 4y ago[deleted]
- jdamon96 4y agoincredible!
- pupe151139 4y ago
- pupe151139 4y ago
- pupe151139 4y agoCom.Samsung.Android.Game.gos:2290:9908:313b42360002
- pupe151139 4y agoCom.Samsung.Android.Game.gos:2290:9908:313b42360002
- gdubs 4y agoI've been playing around w/ StableDiffusion animation using the "deforum" notebook. It's taken a bit to really understand how to get results I like with it, but I'm super happy with how this one came out: https://twitter.com/dreamwieber/status/1565008078466326528?s=21&t=93XvRkPe07uzwpikT4GXaw https://twitter.com/dreamwieber/status/1565008078466326528?s... It's a pretty magical time with this tech. Things are moving very rapidly and I feel excited the way I did when I first rendered 3d animation on my 286 from RadioShack.