4 ms·
> So the (PCI-E) bandwidth strongly affects time to first token On dedicated inference hardware I'd expect model weights to never leave the RAM, and you'd prob
by usrnm 17d ago
> So the (PCI-E) bandwidth strongly affects time to first token
On dedicated inference hardware I'd expect model weights to never leave the RAM, and you'd probably load them on startup before even starting to serve requests