3 ms·
I'd imagine you could exploit something like the stochastic denoising approach from DDIM and descendants, where you add some noise back after each denoising ste
by cheald 2y ago
I'd imagine you could exploit something like the stochastic denoising approach from DDIM and descendants, where you add some noise back after each denoising step, essentially randomly remasking unmasked tokens and giving the model a second chance to unmask them "properly" as the denoised response becomes better-known.