5 ms·
Efficient high-resolution image synthesis with linear diffusion transformer
- smusamashah 2y ago> (e.g. Flux-12B), being 20 times smaller and 100+ times faster in measured throughput. Moreover, Sana-0.6B can be deployed on a 16GB laptop GPU, taking less than 1 second to generate a 1024 × 1024 resolution image.
- deleted 2y ago[deleted]
- echelon 2y agoImage models are going to be widely available. They'll probably be a dime a dozen soon. It's great that an increasing number of models are going open, because these are the ecosystems that will grow. 3D models (sculpts, texture, retopo, etc.) are following a similar trend and trajectory. Open video models are lagging behind by several years. While CogVideo and Pyramid are promising, video models are petabyte scale and so much more costly to build and train. I'm hoping video becomes free and cheap, but it's looking like we might be waiting a while. Major kudos to all of the teams building and training open source models!
- xterminator 2y ago[flagged]
- cube2222 2y agoThis looks like quite a huge breakthrough, unless I'm missing something? ~25x faster performance than Flux-dev, while offering comparable quality in benchmarks. And visually the examples (surely cherry-picked, but still) look great! Especially since with GenAI the best way to get good results is to just generate a large amount of them and pick the best (imo). Performance like this will make that much easier/faster/cheaper. Code is unfortunately "(Coming soon)" for now. Can't wait to play with it!
- liuliu 2y agoIf you read closer to the benchmark, it seems to be slightly worse than FLUX [dev] on prompt adherence and quality. However, the best is to evaluate the result oneself, and the track-record of PixArt Sigma (from the same author?) is pretty good!
- Archit3ch 2y agoIf you generate 25x more images, you can afford to cherry-pick.
- cube2222 2y agoIt would be interesting to have benchmarks that take this into account (maybe they already do or I’m misunderstanding how those benchmarks work). I.e. when comparing quality between two different models of vastly different performance, you could be doing best-of-n in the faster model.
- Vt71fcAqt7 2y agoThat sounds like it could be an intiresting metric. Worth noting that there is a difference between an algorithmic "best of n" selection (via eg. an FID score) vs. manual cherry picking which takes more factors into account such as user preference and also takes time to evaluate, which is what GP was suggesting.
- cube2222 2y agoYeah I’d likely just pick the best scoring one (that is, the pick is made by the evaluation tool, not the model) - to simulate “whatever the receiver deemed best for what they wanted”.
- psb217 2y agoThis is a bit pedantic, but FID score wouldn't really be a viable metric for best of n selection since it's a metric that's only computable for distributions of samples. FID score is also pretty high variance for small sample sizes, so you need a lot of samples to compute a meaningful FID score. Better metrics (assuming goal is text->image) would be some sort of inception score or CLIP-based text matching score. These metrics are computable on single samples.
- lpasselin 2y agoThis comes from the same group as the EfficientViT model. A few months ago, their EfficientViT model was the only modern and small ViT style model I could find that had raw pytorch code available. No dependencies to the shitty framework and libraries that other ViT are using.
- henning 2y ago[flagged]
- ClassyJacket 2y agoThis argument is only fair if you also think human artists should be banned, from birth, from ever looking at any other art. After all that would be training on stolen copyrighted work.
- david-gpu 2y agoDo you believe that human artists should pay license fees for all the art that they have ever seen, studied or drawn inspiration from? Whether graphic artists, writers or what have you.
- kadoban 2y agoHuman artists get in copyright trouble if the spam out a copy of something they studied and sell it. The businesses using AI artists do not seem to.
- ClassyJacket 2y agoImage generation models don't do that either
- david-gpu 2y agoArtists who think that their copyright has been infringed upon are free to sue, just as they do when the alleged plagiarist is a human. I fail to see the difference.
- ben_w 2y agoScale. The cost of the electricity needed to create an image, was the cost of hiring someone on the UN abject poverty threshold to examine it for 10 seconds… with 2 year old models and hardware: https://benwheatley.github.io/blog/2022/10/09-19.33.04.html https://benwheatley.github.io/blog/2022/10/09-19.33.04.html (There's also trademark issues; from the discussions, I think those are what artists actually care about even though they use the word "copyright").
- deleted 2y ago[deleted]
- cpldcpu 2y ago>We introduce a new Autoencoder (AE) that aggressively increases the scaling factor to 32. Compared with AE-F8, our AE-F32 outputs 16× fewer latent tokens, Basically they compress/decompress the images more, which means they need less computation during generation. But on the flip side this should mean less variability. Isn't this more of a design trade-off than an optimization?
- Lerc 2y agoIt might not be compressing more (haven't yet looked at the paper). You can have fewer but larger tokens for the same amount of data. It would decrease the workload by having fewer things to compare against balanced against workload per comparison. For normal N² that makes sense but the page says. We introduce a new linear DiT, replacing vanilla quadratic attention and reducing complexity from O(N²) to O(N) Mix-FFN So not sure what's up there.
- wiradikusuma 2y agoIn my opinion, what's missing in these "image GenAI" tech is the ability to generate subsequent images consistently. That would be useful for e.g. book illustration, comic strips, icon sets. Otherwise, people would think you pick those images all over the internet and not from one source/theme.
- ttul 2y agoThere really are some “free lunches” in generative models. Really impressive work by this group. Ultimately, their model may not be the winner, because so much of what makes a good image gen model is the images and captioning that go into it, and the fine-tuning for aesthetic quality — something Midjourney and Flux both excel at. But the architecture here certainly will get into the hands of the people who can make the next great model. Looking forward to it. This space just keeps getting more interesting.
- amelius 2y agoDoes this finally solve the class of "6 fingers/hand" problems?
- ttul 2y agoThat problem can be fixed through careful fine-tuning, at the cost of losing some generality because the model is punished for drawing bad fingers. This new method outlined in the paper operates in a highly spatially-compressed latent space, but with more channels than previous models, so each latent pixel has 2x the information content than Flux and 8x the content of SDXL. I do wonder whether the high spatial compression means that high resolution features like fingers will be messed up. On the other hand, the higher channel count in the latent space gives the model more detail per pixel to work with… I guess we’ll just have to see.
- cynicalpeace 2y agoNone of this means much to me unless I can actually use it. Sorta like how Sora has been totally overshadowed by Kling, Runway, Minimax. You have to release your model in some fashion for it to be impressive.
- deleted 2y ago[deleted]
- Agentlien 2y agoOn the subject of such high quality video synthesis: have there been any such models which are actually available online? It strikes me that for image synthesis there have been a lot of amazing local models, but I can't remember seeing anything impressive for video which can be run offline.
- cynicalpeace 2y agoThe highest quality ones I mentioned are available via API or web client, but that's enough for me to be happy.
- bick_nyers 2y agoCogVideoX seems to be the best offline model so far