4 ms·
> although time to first token is... not great The token cache works on CUDA too. So yeah, the initial loadout sucks, but almost everything from then on is sol
by dabockster 1y ago
> although time to first token is... not great
The token cache works on CUDA too. So yeah, the initial loadout sucks, but almost everything from then on is solid.
- int_19h 1y agoThis works for scenarios like a simple chat. But the moment you start throwing large documents on it, or lots of code, it's back to waiting.