4 ms·
On my RTX3060 - 12GB VRAM + 24GB RAM , with below command ollama run qwen3.8:27b --verbose "explain mmap”, I got 2.41 Tokens/s, Not sure if that can
by prabhanjana_c 2mo ago
On my RTX3060 - 12GB VRAM + 24GB RAM , with below command
ollama run qwen3.8:27b --verbose "explain mmap”,
I got 2.41 Tokens/s, Not sure if that can be improved considering VRAM doesn’t fit the entire, model.
Additional details:
total duration: 8m18.2870918s
load duration: 612.105ms
prompt eval count: 12 token(s)
prompt eval duration: 2.900965s
prompt eval rate: 4.14 tokens/s
eval count: 1193 token(s)
eval duration: 8m14.660618s
eval rate: 2.41 tokens/s
System spec:
NVIDIA GeForce RTX3060
AMD Ryzen 5 1600 Six-Core
B450 AORUS M Mother board.
NVIDIA-SMI 620.02 Driver:620.02, CUDA Version: 13.2