3 ms·
Rule of thumb I have used is to check the size and if it fits into your GPUs VRAM then it will run nicely. I have not ran into a llama that won't run, but if i
by nextlevelwizard 3y ago
Rule of thumb I have used is to check the size and if it fits into your GPUs VRAM then it will run nicely.
I have not ran into a llama that won't run, but if it doesn't fit into my GPU you have to count seconds per token instead of tokens per second