9 ms·
StoryDiffusion: Long-range image and video generation
- zhoudaquan21 2y agoHi guys, thanks for your interest. The paper and the code are now released: https://github.com/HVision-NKU/StoryDiffusion https://github.com/HVision-NKU/StoryDiffusion. Currently, only the comics-related codes are made public. We are waiting for the company's assessment for the release of the video-related codes.
- cykkkklz 2y agoGitHub Page: https://github.com/HVision-NKU/StoryDiffusion https://github.com/HVision-NKU/StoryDiffusion Paper: https://arxiv.org/abs/2405.01434 https://arxiv.org/abs/2405.01434
- keikobadthebad 2y agoIt'll be good if the girl and the giant squirrel are ever seen in the same park at the same time.
- LeoPanthera 2y agoThe rate of progress of generative AI is honestly quite scary.
- ed_mercer 2y agoReally? Feels like nothing much is happening lately.
- vouaobrasil 2y agoProgress comes in spurts. Due to the negative reactions to AI by some (artists), the system wants it to appear that nothing is happening so that the next wave of AI can be created in relative peace, at which time it will be too late to stop it. We have been conditioned to only react to hype and "news", rather than analyze reality and see the danger.
- thejohnconway 2y agoWhich “system”?
- vouaobrasil 2y agoThe global capitalist system, or the emergent behaviour that comes out of a mass of humanity addicted to technological development through wealth accumulation.
- thejohnconway 2y agoSuch a system can't want anything.
- vouaobrasil 2y agoIt's a term I use for emergent behaviour. And some philosophers of technology would disagree with you, such as the panpsychists. We are just a bag of cells and yet we speak of "wanting" things even though we might just be deterministic bags of blood.
- newswasboring 2y agoWhat are you talking about? ChatGPT-3 came out less than 4 years ago. Stable diffusion's first version around that too. In less than 4 years we went from nothing to making janky but believable video clips. This is not fast enough for you?
- ed_mercer 2y agoI’m just saying that compared to only a few months ago, everything seems to have stagnated. There used to be a lot more news and things getting released left and right.
- freefruit 2y agoSo is Amazon flooded with hyper niche e-books yet?
- selalipop 2y agoI’m working on a platform for reading hyper niche e-books: https://tryspellbound.com https://tryspellbound.com I don’t think this form of generative AI needs to become a source of spam, carefully designed platforms can let people enjoy their niche content without making them feel isolated
- surfingdino 2y agoToo late, it has become a source of spam.
- selalipop 2y agoNot really useful to give up the fight in the infancy of something with as much surface area as generative AI. Is being used to create spam is not the same as needs to be spam, and we mostly just need platforms that leverage generative AI natively to bridge the gap.
- surfingdino 2y agoThere is literally zero need for tools to generate text. Humans generate tons of spam already.
- selalipop 2y agoMy users don't find what these tools generate to be spam. They're enjoying a classic format with a novel level of flexibility and (understandably) find that very fun.
- m463 2y agoI went to buy an air fryer. There were several specific-air-fryer-model recipe books available. But they were all garbage auto-generated stuff. I complained to amazon, and they said since I hadn't purchased the book they couldn't do anything. So I bought the book, complained, and returned it. The chapters devoted to the details of the specific air fryer model were either very general (almost quotes of product description on amazon), or just plain wrong. What I thought I would get would be like the magic lantern books about specific camera models. Instead it was auto-generated pages of nonsense.
- samspenc 2y agoNormally I don't mind spelling errors - and there are plenty in the examples - but my question is, did the system really produce "lunch" when the prompt was "they have launch at restraunt" (verbatim from the sample)? I would imagine it got restaurant right, but I would have expected it to produce something like a rocket launch image instead of figuring out the author meant lunch.
- dkarras 2y agotransformers / attention is very robust against typos as they take the entire context into account just like we do. launch any free LLM and ask them questions with typos that you would notice and auto-correct and you'll see that the models just don't care and understand them. actually they are so resilient that they understand very garbled text without breaking a sweat.
- BoorishBears 2y agoThere's honestly something uncanny about how well they do. In the "early days" of GPT-4 I tried testing it as a way to get around poor transcription for an in-car voice assistant. It managed: "I'm how dew yew say... Freud?" => Turn up the temperature... which was nonsense most people would stare at for a long time before making any sense of.
- noneeeed 2y agoI often use ChatGPT in learning spanish, I find it's great for explaining distinctions between words with similar meanings where a dictionary isn't always a lot of help. I am constantly surprised by how well it copes with my typos, grammatical errors and generally poor spelling.
- neckro23 2y agoAnd if it the model is supposed to be so attentive to context, why did it show a desert instead of "dessert"? After all, they just ate "launch".
- yorwba 2y ago
- hbbio 2y agoGitHub link is not public yet? https://github.com/HVision-NKU/StoryDiffusion https://github.com/HVision-NKU/StoryDiffusion
- stanislavb 2y agoSeems so. I was about to report about it, too.
- smcnally 2y agoThat repo’s not listed https://github.com/orgs/HVision-NKU/repositories https://github.com/orgs/HVision-NKU/repositories
- ActionHank 2y agoA lot of these AI-related announcements seem to be doing this sort of baiting. "I made a new thing", go to the repo, COMING SOON. Or this, here's the paper, no we won't show our work.
- brotherdusk 2y agosorry, i can't access the repo and the pdf doesn't have an href attr, is that by design?
- schoen 2y agoI looked very closely at the videos for a while and managed to find some minor continuity errors (like different numbers of buttons on people's button-down shirts at different times, or different sizes or styles of earrings, or arguably different interpretations of which finger is which in an intermittently-obscured hand). I also think that the cycling woman's shorts appear to cover more of her left leg than her right leg, although that's not physically impossible, and the bear seemingly has a differently-sized canine tooth at different times. But I guess it took me multiple minutes to find these problems, watching each video clip many times, rather than having any of them jump out at me. So, it's not like literally full consistent object persistence, but at a casual viewing it was very persuasive. Maybe people who shoot or edit video frequently would notice some of these problems more quickly, because they're more attuned to looking for continuity problems?
- nyokodo 2y ago> But I guess it took me multiple minutes to find these problems I’m no video editor but I noticed straight away that The characters’ eyes and hair tend to change, sometimes dramatically as they turn their head. Also, the head movement tends to be jerky or abrupt especially in the middle of the turn.
- justinclift 2y agoEyes and teeth seem like they still need further work. Still, looks like things are improving. :)
- godelski 2y agoDid you miss the fish?[0] You should see the error in first viewing What about the woman with glasses? Her face literally "jumps"[1] Same with this guy's hands[2] Interesting, we notice that [1] has "sora" in the name though I think it is a reference to the main image on sora[3] Not sure if the gallery is weird to anyone else, but it doesn't exactly show new images and the position indicator is wonky. The thing that makes me most suspicious is seeing the numbers on these demos. 1, 2, 4 (terrifying to me), 5, 65, 66, 68, 72, 73, 83, 85, 86 (is this Simone Giertz? Vic Michaelis?). The part that is tough about evaluating generative models is the cherry picking for demonstrations. You have to do it or people tear your work apart but also in doing so you give a false impression of what your work can actually do. IMO it has gotten out of hand and is not benefiting anyone. It makes these papers more akin to advertising than communication of research. We talk about integrity of the research community and why we argue over borderline works but come on, if you can get a better review by more samples, you can get better reviews by paying more, not by doing better work. A pay to play system is far worse for the integrity of ML (or any science) than arguing over borderline works. Edit: I think it is also a bit problematic that this is posted BEFORE the arxiv link or GitHub goes live. I'd appeal to the HN community to not upvote these kinds of works until at least the paper is live. [0] https://storydiffusion.github.io/MagicStory_files/longvideo/demo_0065.mp4 https://storydiffusion.github.io/MagicStory_files/longvideo/... [1] https://storydiffusion.github.io/MagicStory_files/longvideo/demo_0027sora.mp4 https://storydiffusion.github.io/MagicStory_files/longvideo/... [2] https://storydiffusion.github.io/MagicStory_files/longvideo/demo_0068.mp4 https://storydiffusion.github.io/MagicStory_files/longvideo/... [3] https://openai.com/sora https://openai.com/sora
- peteradio 2y agoThere is a video of two girls. One girl seems to be sticking out her tongue and then blowing a kiss, but the tongue is appearing again mid-kiss. Very arousing stuff I'll say. Keep up the good work microsft or goggle or whoever made it.
- yard2010 2y agoWorse - bytedance
- topspin 2y agoLove how under "Multiple Characters Generation" the white guy is "A Man," whereas the someone else is "An Asian Man." Reminds me of Daryl Gates and the "normal people" quote, thence patrol cars being called "black and normals."
- fnordpiglet 2y agoA probabilistic regression models behavior will just demonstrate the training data. Don’t hate the player, hate the game.
- topspin 2y agoNo hate for any part of this: it's just amusing.
- forgingahead 2y agoGithub link is broken, and I honestly find it frustrating that the only link to code is the theme source and credits?? Is it really that important to give the static page theme that much real estate instead of actual code release for the project?
- 29athrowaway 2y agoTime for Microsoft Chat 2.0 it seems.
- gbickford 2y agoIt's always disappointing when people publish things to GitHub without the intention of collaborating or sharing.
- pmontra 2y agoThe Moon in the sky seen from the surface of the Moon is wrong? Poetic? Funny? Recursive? A demonstration that these models don't understand anything? Add to the list.
- smusamashah 2y agoThis is unbelievably good. Seems better than Sora even in terms of natural look and motion in videos. The video of two girls talking seems so natural. There are some artifacts but the movement is so natural and clothes and other things around are not continuously changing. I hope it does become open source, which i suspect it won't because it's coming from byte dance.
- cchance 2y agoI don't know if thats true, theirs a massive flicker in the guys hair (the one thats mostly black background and black shirt) half way through it completely loses tracking on his hair and it like snap changes.
- smusamashah 2y agoIf you compare this with current state of openly available video models (assuming this will be open too) this is still a leap. If it is going to be closed like Sora than it's comparable. Sora has different kind of artifacts. These artifacts are an improvement over current state.
- speedgoose 2y agoIs there a video of Will Smith eating spaghetti with this model?
- jerpint 2y agoThe videos look incredible, but a lot of the captions are riddled with grammar/syntax mistakes that seem odd for a model to make of that quality.
- MisterTea 2y agoOne day we won't have 3D engines or GPU's but AI chips that generate the scenes without calculating a single triangle or loading a single texture. We just stream in a scene, IP asset seeds provide the characters, plot and story. But even those can be generated in real-time. Video games, movies, anything will be on demand. No one will act. No one will draw. We will just sit and ask for more. Strange times.
- whamlastxmas 2y agoI had this same realization when Sora came out
- spywaregorilla 2y agoHow is this conceptually different from tracking an embedding for a single character or training a lora on it?
- nephanth 2y agoUm, the github link is a 404, and the paper link links to the webpage itself (— the paper is not on arxiv). Probably they put the website on too fast?
- gtoast 2y agoIts really challenging to think of the positive, constructive uses for this technology without thiking of the myriad, life and societal effecting uses for this. Just interpersonally the use of this technology is heavily weighted towards destruction and deception. I don't know where this ends or where researchers who release this technology think this will go, but I can't imagine its going anywhere good for all of us.