3 ms·
If the same expert is chosen for two consecutive tokens then it’ll act like a 37B model running on the GPU for the second token since it doesn’t need to load th
by hmottestad 2y ago
If the same expert is chosen for two consecutive tokens then it’ll act like a 37B model running on the GPU for the second token since it doesn’t need to load that expert from the main RAM again.