4 ms·
Controlnets are just a clever method of finetuning a network to enable additional conditioning. Nothing about it is specific to latent diffusion, so why would a
by ImprobableTruth 4y ago
Controlnets are just a clever method of finetuning a network to enable additional conditioning. Nothing about it is specific to latent diffusion, so why would another approach like pixel space diffusion not work with it?
Also, I'd hold my horses with calling latent diffusion "out of date". The Imagen paper notes that the results greatly scaled with text-encoder size and the model DeepFloyd (it's a team) are working on uses T5-XXL (the unet is also a decent bit larger, but the Imagen paper claims that this should have a much smaller influence). Latent diffusion might also be able to work with text if you scale up the text encoder.
I'm very skeptical of Google's "trust me, I have the hottest model, it just goes to another school". Muse might have nice fidelity and be fast, but the example images have worse visual quality than Imagen. I definitely wouldn't be surprised if transformers took over, but I don't think it's a foregone conclusion.