3 ms·
> I'd gladly take A5B or A8B or even A10B as a sort of middle ground. Whats up with focusing on the active param count? Do yall fiddle with the weights or some
by Alifatisk 2mo ago
> I'd gladly take A5B or A8B or even A10B as a sort of middle ground.
Whats up with focusing on the active param count? Do yall fiddle with the weights or something?
- martinald 2mo agoYou can run these on CPUs at a somewhat reasonable speed.
- KronisLV 2mo agoOr (somewhat) low TDP GPUs for that matter, like workstation ones, that might have enough total VRAM but not the best bandwidth/compute.
- kennywinker 2mo agoTotal param count decides how much vram you need to run it. Active param count decides how fast it runs. My 10 year old GPU can load quantized 35B or 27B, but it can’t process 27B parameters per token faster than 2-4tok/s, while it can do A3B at >40tok/s
- Alifatisk 2mo agoThank you Kenny