3 ms·
In contrast to memory-bounded autoregressive decoding, diffusion decoding can be computation-bounded. Devices with a decent amount of compute but very small mem
by RandyOrion 1mo ago
In contrast to memory-bounded autoregressive decoding, diffusion decoding can be computation-bounded. Devices with a decent amount of compute but very small memory, e.g. nvidia consumer-grade GPUs, could benefit from diffusion decoding massively.
Performance-wise, I don't think diffusion decoding should be much worse than autoregressive decoding, see https://arxiv.org/html/2604.11035 https://arxiv.org/html/2604.11035 .
As a local LLM user, I really want to see more small diffusion LLM/VLMs with decent performance coming.