3 ms·
True... It would be very interesting to make a comparison of various open models based on token generation speed on these platforms. Presumably starting st some
by oliwary 1y ago
True... It would be very interesting to make a comparison of various open models based on token generation speed on these platforms. Presumably starting st some size the larger accessible RAM wins out over raw speed but low VRAM? Although I suppose things like MoE and FP would also matter.