4 ms·
llama.cpp + Qwen3-4B running on older PC with AMD Radeon GPU (Vulcan). Users connect via web UI. Usually around 30 tokens/sec. Usable.
by lovelydata 11mo ago
llama.cpp + Qwen3-4B running on older PC with AMD Radeon GPU (Vulcan). Users connect via web UI. Usually around 30 tokens/sec. Usable.
- NicoJuicy 11mo agoWhat do they use it for? It's a very small model
- embedding-shape 11mo agoAutocomplete words, I'd wager, as yeah, super tiny model that can barely output coherent output in many cases.