2 ms·
What Jalapeño has done is pair huge bandwidth (almost as much as a full nvidia rubin) with a less powerful processing core. This means that even during the inf
by fancyfredbot 1mo ago
What Jalapeño has done is pair huge bandwidth (almost as much as a full nvidia rubin) with a less powerful processing core.
This means that even during the inference stage, when most architectures are bandwidth bound, they are compute bound instead.
This is why they aren't using multi token prediction, as MTP only helps you when you are bandwidth bound.
The power saving they get is likely from a much lower clock.