7 ms·
SANA-WM, a 2.6B open-source world model for 1-minute 720p video
- mjgil 5mo ago[flagged]
- pferdone 5mo agoWho wrote your comment?
- andai 5mo agoAt this point we should cut our losses and just give Claude an official HN account.
- semiquaver 5mo agoStop posting slop.
- mjgil 5mo agoless security issues with slop
- jaspanglia 5mo ago[flagged]
- rvz 5mo agoGiven that is where everything is going, why not just get there faster by open-sourcing Seedance 2.0, Happyhorse, Veo 3 and all the others.
- carlos-menezes 5mo agoBot comment.
- bobkb 5mo agoThe trouble is the lack of training available to these models compared to the ones like Seedance and Kling who seems to be tapping into their unlimited video inventory. Many models like LTX is technically good but when it comes to slightly different camera movements or the subject interacting with objects they struggle. For a recent example we had to use sample videos generated by closed source models and then use the same for final video.
- vessenes 5mo agoI tend to think of these NV Labs models as architectural demos and ‘free razor blades’ — they’re more intended to inform internal R&D, get customers something that lets them do what they want quickly, and enhance the state of the art. In this case, what looks interesting is the one minute coherence and the massive speedup - they claim 36x over open models with similar capabilities. You can tell they aren’t aiming for state of the art visuals — looks very SD 1.5 in terms of the output quality.
- bobkb 5mo agoAgreed the marketing angle. But beyond the marketing angle what seems to matter is the access to data - look at Seedance , various Kling models etc which are far ahead of others.
- vessenes 5mo agoIt’s hard not to believe that Google doesn’t have an amazing model in-house with all that Youtube content available. But, agreed the Chinese models seem best in the last year or so, and agreed an open policy on training data def makes for better quality
- Fischgericht 5mo agoSo, where is the download? I can't find it on Github, and on your web page the download button is disabled. Also, will this run on RTX 4090 with 24GB memory? Thank you!
- mjgil 5mo agoScroll down and there are more videos --- seems like models will be there "soon".
- ireadmevs 5mo ago> SANA-WM uses only ~213K public video clips with metric-scale pose supervision, completes training in 15 days on 64 H100s, and generates each 60-second clip on a single GPU; its distilled variant runs on a single RTX 5090 with NVFP4 quantization to denoise a 60s 720p clip in 34s.
- beingforthebene 5mo agoThere's a 5 second version available: https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_720p https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_7...
- pferdone 5mo agoFirst video with the guy walking the mountain in snow has consistency issues with the cave entrance. Which is "expected" at this model size?!
- Leonard_of_Q 5mo agoMost videos seem to have some issues like that, e.g. the book on the table in the library video takes up different shapes every now and then. The 'Refiner' effect seems to do the opposite if the examples are representative as in all cases the 1-st stage images look better than the 'refined' ones. Less clutter, more realistic, less 'cowbell' for those who know the phrase.
- andai 5mo agoMy dreams have it too, which is unexpected at that model size!
- ricardobayes 5mo agoThat's actually a great analogy. It's very akin to "generating" dreams.
- echelon 5mo agoRemember the first Will Smith spaghetti?
- agentifysh 5mo agoYeah it got ridiculed and people wrote it off as that was somehow the limit and that wasn't going to change which seems to be the common premise to which people launch their criticisms of AI And those same people forget that its been 3 years since that awful will smith spaghetti video to what we have today which is the beginning of controllable real time videos aka games
- notnullorvoid 5mo agoAll of the videos have rather glaring consistency issues when direction shifts back to areas previously shown.
- mejutoco 5mo agoThey all look like video games. I guess Unreal Engine is used to create synthetic data for training.
- deleted 5mo ago[deleted]
- Royce-CMR 5mo agoThe site notes all starting images are GPT2 imagine gen or nano banana - so that’s probably part of it too.
- joenot443 5mo agoWhat’s the long term utility of world models? There’s no doubt they’re technically impressive, but what does one do with it?
- ACCount37 5mo agoThey can be base models for a bunch of things. Turning text-conditioned video generation models into robotics VLAs is a fun exercise. This one is probably too small to be useful for that, and not diverse enough? But I could be wrong.
- whynotmaybe 5mo agoIt's a step towards something else?
- bix6 5mo agoDigital twin?
- Leonard_of_Q 5mo agoGames. Build campaigns in hours instead of months. Make it possible for users to create their own campaigns, move the action to different game worlds - 'gimme Mario Kart in the ${favourite_game} world', etc.
- AshleyGrant 5mo agoYeah, but is this really that great? Are these models going to remember the town you wandered through on your session yesterday and want to return to? Imagine playing Read Dead Redemption 2 and you attempt to ride your horse from Saint Denis to Valentine and Valentine no longer exists, or is a completely different town located half a mile off from where it was originally. I just don't see how this would work...
- dyauspitr 5mo agoYes, a lot of models don’t state this explicitly, but they can be made deterministic. Not the generation itself, but the same prompt, with a generation seed will always result in the same output.
- mccoyb 5mo agoI struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this intentionality often feel dead in contrast. You run into experiences which break immersion, or pull you out of the experience that the developer is trying to convey to you. It's difficult for me to imagine world models getting to a place where this sort of intentionality is captured. The best frontier LLMs fail to do this in writing (all the time), and even in code, and the surface of experiences for those mediums often feel "smaller" than the user interaction profile of a video game. It's not clear how these world models could be used modularly by humans hoping to develop intentional experiences? I don't know much about their usage (LLMs are somewhat modular: they can produce text, humans can work on it, other LLMs can work on it). Is the same true for the video output here? All this to say, I'm impressed with these world models, but similar to LLMs with writing, it's not really clear what it is that we are building towards? We are able to create less satisfying, less humane experiences faster? Perhaps the most immediate benefit is the ability for robotic systems to simulate actions (by conjuring a world, and imagining the implications). In general, I have the feeling that we are hurtling towards a world with less intentionality behind all the things we experience. Everything becomes impersonal, more noisy, etc.
- deleted 5mo ago[deleted]
- robot_jesus 5mo agoBy and large I agree, but it doesn’t need to be either/or. Many of the most popular games in the past decade are procedurally generated and have nothing “intentionally” placed (apart from tuning/tweaking the balance of the seeding algorithms).
- mccoyb 5mo agoRight, and I wondered how these world models might be use in a careful way (just as agents can be used carefully to accelerate work). Are video game developers using these systems in their workflows? Would love to learn more!
- Incipient 5mo agoOutputting video of that quality/consistency at 1 minute, for a 2.6B model seems insane?
- tredre3 5mo agoIt's because it is insane/misleading. It's a two stage process, scroll to the key features: > A dedicated 17B long-video refiner sharpens texture, motion, and late-window quality on top of the long-rollout backbone.
- basilgohar 5mo agoIt's a very specific use case. This model can generate 1 minute videos of what is essentially a streaming game scene.
- trunkiedozer 5mo agoIt ain’t open source until it’s released. It’s baitware.
- CommanderData 5mo agoAll video models are terrible at consistency. Even closed source ones. Seedance 2.0, Kling 3 are regarded the best closed source video models we have. I have subscribed to a few AI video subreddits, consensus atm is they are good for anything but long form videos with humans. No surprises that we're very good at spotting even the most subtle differences while looking at other people.
- adenta 5mo agowhat subreddits do _you_ subscribe to? I've been doing some content with people at https://industrialallusions.com https://industrialallusions.com
- CommanderData 5mo agohttps://www.reddit.com/r/KlingAI_Videos/ https://www.reddit.com/r/KlingAI_Videos/ https://www.reddit.com/r/HiggsfieldAI/ https://www.reddit.com/r/HiggsfieldAI/ Higgsfield have multiple models available, people use Kling usually 2.5 & 3. There are a few good examples posted right now you'll notice the subtle differences. I have tried to generate things myself and it's extremely hard to have more than 7-8 clips that are consistent, eventually you'll accept a compromise. I think it's why there isn't any long form content being done yet. Getting good results is sometimes just "chance" regardless of how many reference data you have.
- agentifysh 5mo agoRelax its only been 3 years, it's going to get a lot better not worse from here on.
- sebringj 5mo agoi see this and think about Suno's playbook where this could go... survival of the fittest rules the boards where you have user-generated-dynamic video games, not just static ones where design is fixed, the design will be adaptive... based off several prompt input boxes for various things and adhoc while playing, higher tier design boards and the like, this is all going toward user-gen commercial / vanity / personal enjoyment.
- agus4nas 5mo agoIncreíbles resultados
- xyzsparetimexyz 5mo agougly slop
- agus4nas 5mo agoHas anyone actually tested this for robotics simulation? Curious how it handles edge cases in physical environments.
- notnullorvoid 5mo agoJudging by the examples it wouldn't be useful for that, the environments show little physical consistency.
- jubilanti 5mo agoModel weights coming "soon" == currently vaporware. So the weights aren't even open, how can this be "open-source"? Everyone is right to be skeptical of this coming from a 2.8B model. Weights or it didn't happen.
- fc417fc802 5mo agoClearly it's not open in that case. I wonder if we can get the title changed?
- deleted 5mo ago[deleted]
- oersted 5mo agoTo be fair, their whole codebase is open-source, which is better than most open-weight models. But I do agree with the sentiment. https://github.com/NVlabs/Sana https://github.com/NVlabs/Sana
- nl 5mo agoThe model is out here: https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_720p https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_7...
- avaer 5mo agoIs it the same model? The naming is confusing but my understanding is the WM version is not the same as the Video, though I'm sure they are highly related.
- ricardobayes 5mo agoI don't think it's the same - it was uploaded two months ago.
- yieldcrv 5mo agoReally great for visuals during a dj set at a festival or YouTube
- utopiah 5mo agoNice, now instead of just reading slop you'll soon be able to experience slop Worlds, in 3D! /s It's honestly impressive, on the surface. The visuals are gorgeous... but it's still empty. What makes a "World" a world is precisely it's coherency. It's not about how it looks but rather how it "works". The plants in an ecosystems are a certain way because of the available resources, all the way to forces like gravity. It doesn't just "look" like that. To echo Konrad Lorenz a fish doesn't just swim in the water, rather the fish IS an efficient representation of the water it lives within. Here in such "worlds" there is nothing happening. There is minimal superficial coherence, no logic, nothing. The ultimate liminal spaces.
- macwhisperer 5mo agoai is exciting because it shows us what really matters...
- ionwake 5mo agoi survived flash, jquery, svn, soap, xml, microservices and crypto now some norwegian teenager is generating netflix-quality worlds during lunch break from a jpeg of a forest EDIT> dont ask how I came up with this quote
- PyWoody 5mo agoI tried watching the cave video and I was immediately overcome with nausea. I've never experienced anything like that before in my life. Wild. I can't say I'm looking forward to an AI video future.
- rpozarickij 5mo agoWhen I installed very high quality (CRI 98, R9 94, virtually flicker free) light bulbs to one of my apartment rooms I had headaches and felt occasional confusion for about a week while being in that room, so I had to slowly increase the amount of time the light bulbs were turned on for. To my understanding my brain was very used to the way objects and lighting looked in that particular room so it needed to rewire some knowledge given that I've spent many thousands of hours in that room with previous light bulbs. I'm curious if a younger me would have adapted much faster.
- resist_futility 5mo agowarning: viewing the videos that auto play on that page shot up my downloads to 350Mbps on that page
- harshreality 5mo agoI only noticed after more than an hour with the page left open in a tab. Is it really streaming and re-streaming the same videos? There's too much to cache so it keeps re-transferring them indefinitely? I hope nobody leaves that page open on a metered or capped network connection. I'm surprised github hasn't suspended the page. Are AI researchers so used to burning through compute and network resources that they don't stop to think about a webpage that will autoplay and loop multiple HD videos?
- debugnik 5mo agoNearly every website for papers about AI applied to graphics hangs my phone browser, so I'm assuming the answer's yes.
- manquer 5mo agoThey don't even notice it happening, it is not a conscious thought not to fix it. Empathizing about problems you don't face is a hard product/ux and management skill. Facebook famously simulated 2G on Tuesdays 10 years ago[1] for example to get their employees to see the problems their users have.[2] People don't to put effort in noticing(solving comes next) problems they don't face. It is why things like a11y and i18n need regulation like ADA etc. [1] https://engineering.fb.com/2015/10/27/networking-traffic/building-for-emerging-markets-the-story-behind-2g-tuesdays/ https://engineering.fb.com/2015/10/27/networking-traffic/bui... [2]While it would be hard to attribute directly, GraphQL and to an extent React probably was influenced by these kind of things
- ssl-3 5mo agoIt's really something. It was using ~400Mbps of my mostly-idle 500Mbps connection (slow for some folks, but pretty speedy in my particular ghetto). It appears that there are 62 videos on the page. They're generally 16fps and 60s long. All are h.264, 1280x704. The median bitrate is 4.962 Mbps. I don't know enough about JS to try to understand WTF it is doing, but there's only 1.3 GB of video on that page. At a transfer speed of 400Mbps, the whole mess of them should be downloadable in around 30 seconds. But it wasn't behaving that way at all. It instead behaved as an excellent bandwidth-waster. (Woe to those who click this link on metered connection, I guess.)
- deleted 5mo ago[deleted]
- alloyed 5mo agosilly question: what's "world" about what's being generated here? is the an actual abstract representation of physical space (like, eg, a game-engine style scene graph?) or does it just mean "this video generator is more coherent physically than other video generators"
- futureshock 5mo agoWorld in this context means that these videos are interactive, just like a video game. In the linked examples you can see the keyboard and mouse inputs. The model is trained to maintain about a minute of scene consistency so you can look around and objects out of view will reappear when you look back in that direction.
- oersted 5mo agoA world-model is one that predicts the next state of a simulated world given the current state and optionally some action from an agent inhabiting the world. It is quite analogous to a language-model that predicts the next word. That world-state can be anything, but in the last year or two, the term has taken a narrower meaning: a video generation model that reacts naturally to game-like controls, as if it was simulating a videogame. But there's no additional state behind the video frames.
- agentifysh 5mo agoRunning this on GPU is quite impressive. I see some people expressing discontent and worries but we are early and this is the worst its going to be, I am very excited to see the impact this will have on games
- bilsbie 5mo agoWhat would it take to get this on VR? Anyone looking into it?
- futureshock 5mo agoIt is plausible, the model would just need to be trained on a lot of stereoscopic data.
- maxignol 5mo agoI can’t seem to grasp why everyone says only slop gets produced by AI models (and particularly those world models). Imo it’s shit in -> shit out. Great work can be achieved using those. Slop gets produced by careless users.
- mkl 5mo ago2.6B, but then: > A dedicated 17B long-video refiner sharpens texture, motion, and late-window quality on top of the long-rollout backbone.
- w10-1 5mo agoGist: > 720p, 1-min video generation with 6-DoF camera control As nl said, > The model is out here: https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_720p https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_7... README says "intended for research use only" Code license is Apache 2.0 Model license (nvidia open...) says Models are commercially usable. You are free to create and distribute Derivative Models (As usual: model output is unrestricted, and also unprotectable absent human authoring)
- avaer 5mo agoThe linked model does not claim to support camera control, it doesn't appear to be SANA-WM. You might be getting confused.
- deleted 5mo ago[deleted]
- ArchieScrivener 5mo agoThe samples are creepy and lonely. Zero group/crowd or face images, just lonely empty perspective. Very hollow feeling.
- ricardobayes 5mo agoYes because most spaces shown are liminal, giving that eery feeling like an empty shopping mall. Still an interesting concept though.