2 ms·
LLM prompt processing and diffusion models are compute bound, while LLM token generation is memory bandwidth bound.
by MrScruff 2mo ago
LLM prompt processing and diffusion models are compute bound, while LLM token generation is memory bandwidth bound.