3 ms·
I would expect that 4060ti to get about 20-25 tokens per second on Mixtral. I can read at roughly 10-15 tokens per second so above that is where I see diminishi
by pitched 3y ago
I would expect that 4060ti to get about 20-25 tokens per second on Mixtral. I can read at roughly 10-15 tokens per second so above that is where I see diminishing returns for a chatbot. Generating whole blog articles might have you sit waiting for a minute or so though.
- thrdbndndn 3y agoThanks, that sounds more than tolerable than "more than an hour"! I also have the 16GB version, which I assume would be a little bit better.
- cellis 3y agoIt depends on the context window, but my 3090 gets ~60/s on smaller windows.
- visarga 3y agoI get 50-60t/s on Mistral 7B on 2080 Ti