3 ms·
A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that make
by Eisenstein 12d ago
A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.
- beastman82 12d agoOff the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
- deleted 12d ago[deleted]
- medvezhenok 12d agoI think you’re missing that MTP can predict more than 1 token in advance.