8 ms·
FastVideo: a lightweight framework for accelerating large video diffusion models
- zhisbug 2y agoHugginface model and data link: https://huggingface.co/FastVideo https://huggingface.co/FastVideo
- Reubend 2y agoDoes the distillation done here have a large impact on quality compared to the original "slow" models?
- fc417fc802 2y agohttps://huggingface.co/FastVideo/FastHunyuan#evaluation https://huggingface.co/FastVideo/FastHunyuan#evaluation Seems like the images are sharper and noisier. The originals seem to have more blur.
- woodson 2y agoTo be fair, one would use more than 6 steps with the original Hunyuan model, so perhaps that’s why they’re so blurry. But that’s even slower.
- deleted 2y ago[deleted]
- holoduke 2y agoWe need videocards with lots of memory. Give me a 4080 with 192gb. I would be happy. We really need AMD to come up with new cards to wake up NVidia and start some fierce competition
- jsheard 2y agoIt's not really feasible to scale GDDR-based designs that big. The 5090 is expected to have 32GB which probably means the workstation variant will have 64GB, but that's the limit of the conventional GPU memory architecture for now. HBM is fast and high capacity but prohibitively expensive, and LPDDR is cheap and high capacity but relatively slow, so there's no free lunch to be had.
- andybak 2y agoWhat would it take to have a unified memory architecture to rival Apple's ? Is it theoretically possible with PC motherboards and GPUs that sit in card slots of some form?
- sroussey 2y agoNVIDIA Jetson?
- deleted 2y ago[deleted]
- girvo 2y ago> Is it theoretically possible with PC motherboards and GPUs that sit in card slots of some form? As far as I'm aware, not with the speed that unified memory gets. The trace and path lengths alone put hard caps on what can be done in terms of signalling speed. But I'm not an expert, I'm recounting what I was told when I asked the same question! Perhaps the state of the art has improved in this space?
- treesciencebot 2y agoFor anyone that wants to test the original (non-distilled) HunyuanVideo (which is an amazing model) we have 580p version taking under a minute and 720p version taking around 2.5-3 minutes in our playground: https://fal.ai/models/fal-ai/hunyuan-video https://fal.ai/models/fal-ai/hunyuan-video (it requires github login & and is pay-per-use but new accounts get some free credits).
- echelon 2y agoOpen source video models are going to beat closed source. Ecosystem and tools matter. Midjourney has name recognition, but nobody talks about Dall-E anymore. The same will happen to Sora. Flux and Stable Diffusion won images, and Hunyuan and similar will win video. Hunyuan, LTX-1, Mochi-1, and all the other open models from non-leading foundation model companies will eventually leapfrog Sora and Veo. Because you can program against them and run them locally or in your own cloud. You can fine tune them to do whatever you want. You can build audio reactive models, controllable models, interactive art walls, you name it. Sora and Veo just aren't interesting. They're at one end of the quality spectrum, and open models will quickly close that gap and then some.
- creato 2y agoI'm curious what your take on GIMP vs. Photoshop would be?
- echelon 2y agoNobody is itching to put GIMP into their product, but everyone can think of ways to build upon Llama and Flux and provide new value.
- whywhywhywhy 2y agoIt’s not comparable because GIMP has never had the effort put into it to compete with Photoshops most basic features. 15-20 years ago they were arguing that adjustment layers were not needed and they only managed to ship some form of it this year. Blender vs commercial 3D software is a better example.
- MarthaIrene528 2y ago[dead]
- pauloday 2y agoSomeone wrote the following comment then deleted it. I spent 30 minutes on my response and wanted to post it anyway. Apologies if the original comment was deleted by a mod, I hope this is OK to post. ---QUOTE--- My "test" for video generation turning movie making on its head is when a model can add the missing Tom Bombadil chapters to Peter Jackson's LOTR movies. Probably 20 - 30 minutes of HD, aesthetically synced, scripted etc with minimal editing after a detailed prompt and source material. Qualifier - the AI just has to follow the book script, third party tools ok to use for lip syncing and audio :) I said 5 years away last year. Feels like it might be more like 1 - 2 years. What do you think? ---END QUOTE--- My response: I think we're getting into diminishing returns territory with this AI stuff. These video/image generators are impressive but they don't "understand" physical reality and probably never will without a breakthrough. You can see this in the demo videos, the best looking ones are glorified still images and the worst are whenever something physical happens, like the lemon being picked up or the guy eating cereal. These examples may get better, but I really doubt they'll ever look like real unaltered camera footage without adding an understanding of how our physical reality works into the model somehow. For the script generation, Fellowship of the Ring is not a movie script and requires serious interpretation and planning to be converted into one. Especially if you want it to fit into Jackson's films at all. If nothing else the dialog and frequency of songs/poetry are very different. The current text generators aren't really capable of that kind of planning yet, but I wouldn't be surprised if there's a screenwritten treatment of that chapter floating around on the internet somewhere, or at least bits of one. It has certainly ingested The Fellowship of the Rings, and plenty of screenplays plus the books they were based on. So maybe chatgpt can make a convincing script. I asked the free version and got some dialog that seems fine, but absolutely no scene direction at all. I'm willing to believe that was either an issue with my prompting or something that can be fixed in 5 years. So at least the script may be possible. As for converting it into an actual piece of film, I don't think that's currently possible without a breakthrough on planning. There's a reason these video demos aren't usually very long, it's because they aren't good at scene changes. People's faces change, rooms change shape, etc. Maybe that can be fixed through engineering, but film editing is hard. It's not easy to plan and chain together shots in a way that gives a proper sense of physical reality while conveying everything a scene needs to. Take a look at Dan Olsen's video analyzing the editing of Suicide Squad[1]. That movie was edited by a trailerhouse and it shows. A big issue is that the scenes and shots don't flow together very well - it's edited like a bunch of separate shots and scenes rather than a coherent whole. As a result it's generally considered one of the worst films big budget ever made. And from my (admittedly limited) understanding/playing around with these generators, they aren't even remotely close to being able to do the type of planning needed to pull that kind of editing off, much less something on the level of Jackson's adaptation. Again I could be wrong but it really seems like another "Attention is All You Need" level breakthrough to get there. So I'd say no, I don't think we'll get what you describe, at least not at any level of quality, in 1-2 years. 5 years sounds more realistic but I really believe we'd need another huge breakthrough to get there, and those are hard to come by. Assuming one will happen in any given time period seems foolish. But a lot of smart people are working on that, so maybe we'll get it. But I don't think we'll even get there in 10 years with just engineering improvements on the current stuff. Scientific progress isn't linear. Yours and a lot of other predictions about AI stuff really remind me of how all the futurists in the 50's thought we'd be able to freeze and unfreeze humans in a few short years. They thought that because it's actually really easy to do that with hamsters, but it turns out scaling the process up isn't so easy (Tom Scott has a good video tangentially related to this[2]). I think a lot of people are standing near the top of the steep part of a sigmoid curve and saying "Wow look how far we've come in just 3 years! The next 3 years are going to be insane!" When in reality we just have a long plateau of minor improvements in front of us. But who knows, maybe that next breakthrough is right around the corner. [1]: https://www.youtube.com/watch?v=mDclQowcE9I https://www.youtube.com/watch?v=mDclQowcE9I [2]: https://www.youtube.com/watch?v=2tdiKTSdE9Y https://www.youtube.com/watch?v=2tdiKTSdE9Y