4 ms·
Understanding classifier-free guidance for style transfer
- lewq 4y agoI thought this was cool because it's very different from the usual anime crap And you can see how smoothly varying the math produces these varying images in response to the prompt
- gyre007 4y agoIt's pretty wild how much value Automatic1111 webuid enabled to create.
- binocarlos 4y agoThis is a really cool blogpost - that you can influence a model with guidance towards a desired outcome feels like a really valuable feature. I'm currently working on a project that is generating stable diffusion images and this kind of technique would be so useful ;-)
- hodesdon 4y agoI wrote this post because I wanted to know what effect one of the main latent diffusion model parameters -- the classifier guidance scale (cfg_scale) -- was having on the sampling process. As well as smoothly varying the cfg_scale, I think it would be fascinating to do some mechanistic interpretability on latent diffusion models. Something like the microscope tool OpenAI used to have for convnets: https://microscope.openai.com/models https://microscope.openai.com/models
- tundoz 4y agoI'm curious what initial results you got when training only on 6 images. Were generated images not Wildsmith-ish enough?
- hodesdon 4y agoIt's hard to be precise, because I don't have a good metric for Wildsmith-ness, but yes, the fewer the images used, the fewer aspects of the style that can be picked up with the embedding. I think the textual inversion authors' claim that you can use just 6 or so images wasn't made with style transfer in mind. Six is probably sufficient when you want to introduce a word for a specific object.