6 ms·
“If you’ve been following the red-hot AI content generation scene, then you’re probably already aware that it’s likely a matter of months before authentic-looki
by tomComb 4y ago
“If you’ve been following the red-hot AI content generation scene, then you’re probably already aware that it’s likely a matter of months before authentic-looking video of newly invented Seinfeld bits like this can be generated by anyone with a few basic, not-especially-technical skills.”
A matter of months eh?
To me that claim looks to be about driving up business. I realize lots of people and companies are making similar claims, and I have no doubt we will get there in a reasonable amount if time, but I believe that people with your level of expertise believe it will be at least at year until we achieve what you decide.
This whole thing is remarkable but now people are overselling it IMO.
- mholm 4y agoI’m of similar opinion. I keep seeing claims of smooth AI video at the level of existing image generation ‘coming soon’ but I haven’t seen any evidence that it’s even close. Few years? Sure.
- doopy1 4y agoI think we can have it very quickly if someone threw a ton of money at it. The question is who wants to take that bet and run at a loss for ages.
- seibelj 4y agoSelf driving cars followed the same pattern - the demo is decent but the 99.9999% case necessary for real-world usage is perpetually out of reach
- metadat 4y agoWhat is acceptable for image and content generation has almost no relation to what is acceptable for safety-critical systems.
- seibelj 4y agoGiven how insanely high the bar is in creative industries, I’m not sure how deep a purely AI generated graphic can succeed
- ChadNauseam 4y agoI don’t think the bar is uniformly that high. Go look at the covers of some Kindle Unlimited titles – Midjourney art is better
- mkaic 4y agoI expect short video generation to be at the level of DALLE 2 in 1 year — e.g. able to get the gist of a prompt, but with lots of artifacts, requiring a lot of compute, and frequently ignoring large parts of the prompt.
- deleted 4y ago[deleted]
- nodja 4y agoDepending on how you're measuring it, video generation will come in less than 12 months, or it's already here. - In the Video Diffusion paper[1] they generate 64x64 16 frame videos, model is not released, but an open source implementation using an imagen-like pipeline exists[2], no model is available publicly. - The CogVideo paper[3], which is essentially a huge transformer model, can generate 480x480 videos, the code for the paper and models are open source[4], but be warned that you need a huge GPU like an A100 to run the damned thing. The future of text to video generation will probably be video diffusion, i.e. using 3D UNets, or more likely a much more optimized version of a UNet. Improvements on diffusion models are happening on a daily basis, and we're probably a couple papers away from having a breakthrough in efficiency that we can generate good looking videos on a high end GPU. [1] https://arxiv.org/abs/2204.03458 https://arxiv.org/abs/2204.03458 [2] https://github.com/lucidrains/imagen-pytorch/tree/main/imagen_pytorch/imagen_video https://github.com/lucidrains/imagen-pytorch/tree/main/image... [3] https://arxiv.org/abs/2205.15868 https://arxiv.org/abs/2205.15868 [4] https://github.com/THUDM/CogVideo https://github.com/THUDM/CogVideo
- nl 4y agoIt's very likely that realistic looking (to the level of Stable Diffusion) video will happen and tools to create it will be available within 12 months (maybe 80% likelihood). What is likely to be missing is the ability to control that video in useful ways directly from prompts. There will be some type of direction, but as anyone who has spent time doing prompt-based image generation actual control isn't there yet.
- nodja 4y ago> but as anyone who has spent time doing prompt-based image generation actual control isn't there yet It's not, but there are many efforts to fix this and they solve most problems, although I admit they're not ideal. One way you can control generation is simply by messing around with per token weights, for example if your prompt is being partially ignored, you can use weights to have the guidance give more emphasis on that part of the prompt. Another way you can control it is by using img2img and input as simple drawing, this can help it better understand say, which color you want things to be. The best tool of all is of course, copious amounts of inpainting and doing composition, eyes not the right color? inpaint them in, messed up hand? inpaint, etc. If your issue is that you can't generate a specific character or style at all, you have tools like text-inversion that can create a special token that represents the rough idea, assuming you already have a couple images that represent that idea already. There is a real fix tho, the big issue with these prompts is that they use the clip text encoder, the clip text encoder was trained only on image captions, which means that it's understanding of the world is limited to whatever is represented by captions that exist on images found on the internet, a very limited subset of language, limiting the quality of the generated embeddings as the model not only is bad at basic language, it doesn't properly understand the relationships between words and such. Image models that use a large language model (LLM) listen to the prompt much better, since these LLM models generate much less noisy embeddings, a LLM contains embeddings big enough that the diffusion model can actually spell, i.e. if you ask for a sign that says something, the sign will come out with proper spelling and font choice. Sadly there's no such model open to the public yet, but stability.ai is currently training such a model and we can expect its release hopefully before christmas. Having a image model trained with a LLM text encoder and then applying the tricks used to improve current models will give you an unprecedented level of control over image generation, the next step is probably having an instruction based model that takes both an image and text as input, and runs the instruction over the image. Example, you give it a image of a person and you send the instruction of "draw this in the style of pixel art" and it'll do it for you, it's already currently possible with img2img, but having a model than listens to instructions can extend this to more useful things like "remove the background", "add another cat", "tilt the sign 90 degrees", "make everything black and white except her dress", etc.
- godelski 4y agoSo we can already make video. But the "authentic-looking" part is going to be quite a bit. Even current face work still has a lot of issues that are a bit more subtle. The diversity in a whole scene is higher. Then also you'd need convincing voice deep fakes, which does exist. But then you need to integrate this and/or the lip syncing into the prompts (totally doable). Right now peopel are trying to smooth variance between frames. I'd suspect we'll get something 360p-480p quality in <2 years (definitely less than 5). Idk, is 24 months "a matter of months"?
- thomasahle 4y ago>> it’s likely a matter of months > A matter of months eh? Right. As always people overestimate the effect of a technology in the short run and underestimate the effect in the long run. It was 15 months between Dall-E and Dall-E 2. And it's already been 5 months since the release of Dall-E 2. To get homemade Seinfeld bits we'll presumably have to wait for some big research org to release a paper and then it's a matter of months on top of that. Unless hobbyists can take the current diffusion technology from images and simple videos to complicated videos and video + sound. But I doubt that.
- mudrockbestgirl 4y ago> but I believe that people with your level of expertise What expertise? It's just a marketing article written by a journalist who writes about all kinds of stuff. It's not written by someone with actual expertise in building models or doing research in the field. I'm surprised this marketing piece is even being upvoted because it's void of anything technical or deep.
- rco8786 4y agoThis seems a bit pedantic? “A matter of months!? Psh, more like a year” Either is insanely fast and the distinction is not particularly notable. Also a matter of months seems an entirely reasonable estimate given the insane velocity around this stuff. I’ve already seen some video generation models.
- tomComb 4y agoI'll be pedantic again ... you changed what I wrote from "at least" a year to "more like" a year, which is actually quite different.