3 ms·
Diffusion Finetuning Myself
- throwaway314155 1y agoYour concern about catastrophic forgetting is mostly unfounded in the regime of fine-tuning large diffusion models. The weights in this case will maybe suffer from some damage to accuracy on some downstream tasks. In general though, it is not “catastrophic”. I believe this is due to the attention mechanism but I’m happy to be corrected.
- frotaur 1y agoI see, it was probably my high learning rate that caused problems. To be honest, I got a bit lazy to retry full finetuning since LoRA worked so well, but maybe I'll revisit this in the future, maybe with Qwen Image.
- throwaway314155 1y agoPerhaps what you were dealing with was actually exploding gradients using fp16 training which _are_ prone to corrupting a model and this can depend on the learning rate.
- adzm 1y agoMinor observation: the formula text appears to go above the sticky header in the website.
- frotaur 1y agoTrue, I hadn't noticed, thanks! I'll try to fix that in the near future. Edit : Fixed!