8 ms·
Stable Diffusion 3: Research Paper
- whywhywhywhy 3y agoIt's impressive that it spell words correctly and lay them out but the issue I have is the text always has this distinctively overly fried look to it. The color of the text is always ramped up to a single value which when placed into a high fidelity image gives the impression of just slapping some text on top with photoshop afterwards in quite an amateurish fashion rather than text properly integrated into an image.
- imiric 3y agoBut the sample images they show here showcase a good job at blending text with the rest of the image, using the correct art style, composition, shading and perspective. It seems like an improvement, no?
- blehn 3y agoThe blending looks better, but the LED sign on the bus for example looks almost like handwritten lettering... the letters are all different heights and widths. Not even close to realistic. There's a lot of nuance that goes into getting these things right. It seems like it'll be stuck in an uncanny valley for a long time.
- viraptor 3y agoIt's just the presented examples. See the first preview for examples of properly integrated text https://stability.ai/news/stable-diffusion-3 https://stability.ai/news/stable-diffusion-3 Especially the side of the bus.
- whywhywhywhy 3y ago"Go" "Stable Diffusion" are not they look slapped on and would be sent back in most studios. Dream on is a bit better but looks very messy, badly spaced, inconsistent letter sizes.
- MyFirstSass 3y agoSide of the bus still looks weird though. Like some lower resolution layer was transformed to the side, also still to bright. Still impressive we've come to this but we're now in the uncanny valley which is a badge of honour tbh.
- bsenftner 3y agoI'm expecting at some point the stable diffusion community of developers to recognize the value of Layered Diffusion, the method of generating elements with transparent backgrounds, and transitioning to outputs that are layered images one may access and tweak independently. The addition of that would make the hands-on media producers of the world say "okay, now we're talking, finally directly indigestible into our existing production pipelines."
- cthalupa 3y agoThere's already ComfyUI nodes for Layered Diffusion. https://github.com/huchenlei/ComfyUI-layerdiffuse https://github.com/huchenlei/ComfyUI-layerdiffuse Of the people I know in the CG industries using SD in any sort of pipeline, they're all using Comfy because a node based workflow is what they're used to from things like Substance Designer, Houdini, Nuke, Blender, etc.
- bsenftner 3y agoThat's my impression as well. I used to work in VFX, my entire career is pretty much 3D something or other, over and over.
- GaggiX 3y agoIt's very likely an artifact of CFG (classifier-free guidance), hopefully some days will be able to ditch this kinda dubious trick. This is also the reason why the generated images have this characteristic high contrast and saturation. Better models usually need to rely less on CFG to generate coherent images because they fit the training distribution better.
- finnjohnsen2 3y agoQuestion is, will SD3 be downloadable? I downloaded and run the early SD locally and it is really great. Or did we lose Stable Diffusion to SAAS also? Like we did on many of the LLMs which started of so promising as for self hosting goes
- sen 3y agoSounds like it’ll be downloadable. FTA: > In early, unoptimized inference tests on consumer hardware our largest SD3 model with 8B parameters fits into the 24GB VRAM of a RTX 4090 and takes 34 seconds to generate an image of resolution 1024x1024 when using 50 sampling steps. Additionally, there will be multiple variations of Stable Diffusion 3 during the initial release, ranging from 800m to 8B parameter models to further eliminate hardware barriers.
- finnjohnsen2 3y agoThanks for pointing that out. Super promising.
- nuz 3y agoThe 800m model is super exciting
- jncfhnb 3y agoIt will probably suck. These models aren’t quite good enough for most tasks (other than toy fun exploration). They’re close in the sense that you can get there with a lot of work and dice rolling. But I would be pessimistic about a smaller model actually getting you where you want.
- cooper_ganglia 3y agoYeah, SDXL is probably better than SD3 800M if I had to guess. I’m looking forward to the quality advancements with LCM Loras, or an SD3 Turbo!
- 3y ago
- TheAceOfHearts 3y agoIt's very exciting to see that image generators are finally figuring out spelling. When DALL-E 3 (?) came out they hyped up spelling capabilities but when I tried it with Bing it was incredibly inconsistent. I'd love to read a less technical writeup explaining the challenges faced and why it took so long to figure out spelling. Scrolling through the paper is a bit overwhelming and it goes beyond my current understanding of the topic. Does anyone know if it would be possible to eventually take older generated images with garbled up text + their prompt and have SD3 clean it up or fix the text issues?
- vergessenmir 3y agoI would imagine with an img2img workflow it would be. The same way you can reconstruct a badly rendered face but doing a second pass on the affected region
- declaredapple 3y agoThe best way to do it right now is controlnets. I'm not sure about re-doing the text of old images - you could try img2img but coherence is an issue, more controlnets might help
- emadm 3y agoYes, this is possible, we have ComfyUI workflows for this.
- mise1 3y agoIt's a surprisingly difficult task with quite a bit of research history. We looked at different solutions extensively (https://medium.com/towards-data-science/editing-text-in-images-with-ai-03dee75d8b9c https://medium.com/towards-data-science/editing-text-in-imag...) and ended up building a tool to solve the problem: https://www.producthunt.com/posts/textify-2 https://www.producthunt.com/posts/textify-2 Eventually models will get to the point where they can do this well natively but for now the best we can do is a post-processing step.
- nodja 3y ago> I'd love to read a less technical writeup explaining the challenges faced and why it took so long to figure out spelling. Scrolling through the paper is a bit overwhelming and it goes beyond my current understanding of the topic. I'm not an ML researcher but I can answer this. Note that this is not information from the paper, just my own findings from following the "scene". We actually figured out spelling not long after diffusion models came out, the imagen paper that came several months before SD1 explained how they did it, which is the same technique SD3 and Dalle3 use, instead of using CLIP's text encoder, they use T5. The reason why image models can't spell is the same reason why language models also have difficulty spelling. Tokenization. Simply speaking instead of seeing each letter individually, we split a sentence into sub-words, most commonly called tokens and that's what the model sees. The model never gets to see each letter individually. But it turns out that if you make the model big enough and feed it enough data, it actually learns how each token is spelled out. Clip is both small (200-500M params) and trained on limited text data (only image captions). T5 is trained on a large corpus of data and is also huge (~5.5B params). This makes T5 the obvious choice if you care about spelling. So why use CLIP in the first place? Simple: it's much easier to train on clip embeddings than T5 embeddings, not only are they smaller, but because clip is trained on text/image pairs and due to backpropagation, the text embeddings also contain a lot of visual semantic information. This simplifies a lot of the work the text->diffusion attention modules need to do. Another reason is that T5 is absolutely massive, it's 5x larger than the image part of the model and 10-20x larger than CLIP. If you take a close look at diagram (a) in the paper you'll actually see that SD3 uses both CLIP and T5, not only that but they trained it in a way that makes the encoder used optional, so you can use the CLIP models only if you don't care about spelling and image composition (CLIP is also bad at understanding prompts), which is useful because most GPUs can't handle T5 on it's own. Tho I suspect someone will distil T5 so it becomes 5-10 times smaller than it currently is at a minimal loss on how good it is for prompting.
- WiSaGaN 3y agoMore and more companies that were once devoted to being 'open', or were previously open, are now becoming increasingly closed. I appreciate Stability AI releases these research papers.
- loudmax 3y agoIt's hard to build a business on "open". I'm not sure what Stability AI's long term direction will be, but I hope they do figure out a way to become profitable while creating these free models.
- sharmajai 3y agoMaybe not everything should be about business.
- TehCorwiz 3y agoI agree, but Y-Combinator literally only exists to squeeze the most bizness out of young smart people. That's why you're not seeing so much agreement.
- phkahler 3y ago>> but Y-Combinator literally only exists to squeeze the most bizness out of young smart people. YC started out with the intent to give young smart people a shot at starting a business. IMHO it has shifted significantly over the years to more what you say. We see ads now seeking a "founding engineer" for YC startups, but it used to be the founders were engineers.
- te_chris 3y agoSqueezed all the alpha out of the idealists now it’s the business guys turn
- bufferoverflow 3y agoIf you agree, do you mind paying a few hundred thousand for my neural net training expenses?
- edshiro 3y agoThis is really exciting to see. I applaud Stability AI's commitment to open source and hope they can operate for as long as possible. There was one thing I was curious about... I skimmed through the executive summary of the paper but couldn't find it. Does Stable Diffusion 3 still use CLIP from Open AI for tokenization and text embeddings? I would naively assume that they would try to improve on this part of the model's architecture to improve adherence to text and image prompts.
- MrCheeze 3y agoOne of the diagrams says they're using CLIP-G/14 and CLIP-L/14, which are the names of two OpenCLIP models - meaning they're not using OpenAI's CLIP.
- deleted 3y ago[deleted]
- MrCheeze 3y agoI have just been informed that my above comment is false, the CLIP-L is in fact referring to OpenAI's, despite that also being the name of an OpenCLIP model.
- ollin 3y agoThey use three text encoders to encode the caption: 1. CLIP-G/14 (OpenCLIP) 2. CLIP-L/14 (OpenAI) 3. T5-v1.1-XXL (Google) They randomly disable encoders during training, so that when generating images SD3 can use any subset of the 3 encoders. They find that using T5 XXL is important only when generating images from prompts with "either highly detailed descriptions of a scene or larger amounts of written text".
- vessenes 3y agoThis looks great, very exciting. The paper is not a lot more detailed than the blog. The main Thing about the paper is they have an architecture that can include more expressive text encoders (t5-xxl here), they show this helps with complex scenes, and it seems clear they haven’t maxed out this stack in terms of training. So, expect sd3.1 to be better than this, and expect 4 to be able to work with video through adding even more front end encoding. Exciting!
- nojvek 3y agoHe! in contrast to Stability AI, Open AI is the least closed AI lab. Even Deep Mind publishes more papers. I wonder if anyone in Open AI openly says it "We're in for the money!" The recent letter by SamA regarding Elon's trial had as much truth as Putin saying they are invading Ukraine for de-nazification.
- edwcross 3y agoNice improvements in text rendering, but it seems generating hands and fingers is still difficult for SD3. None of the pictures in the example contain human hands, except for the pixelized wizard; and the monkey hands seem a bit odd.
- astrange 3y agoThe proper solution to fine details like hands will be conditioning the image on a 3D pose eg with controlnets. It's hard to get exactly what you want with only a single text prompt.
- liuliu 3y agoThis arch seems to be flexible enough to extends to video easily. Hopefully what we have here will be another "foundation" blocks like the transformer blocks in LLaMA. Why: It looks generic enough to incorporated text encoding / timestep condition into the block in all the imaginable ways (rather than in limited ways in SDXL / SD v1, or Stable Cascade). I don't think there is much left to be done there other than to play with positional encoding (2D RoPE?). Great job! Now let's just scale up the transformers and focus on quantization / optimizations to run this stack properly everywhere :)
- tmabraham 3y agoThe paper has preliminary results for video as well