2 ms·
Just waiting on llama.cpp support :) I usually use GPT-oss-120B with CPU MoE offloading. It writes at about 10tps, which is useful enough for the limited thin
by binary132 1y ago
Just waiting on llama.cpp support :)
I usually use GPT-oss-120B with CPU MoE offloading. It writes at about 10tps, which is useful enough for the limited things I use it for. But I’m curious how Q3 Next will work (or whether I’ll be able to offload and run it with GPU acceleration at all.)
(4090)