3 ms·
you can run Qwen 3.6 35B-A3B (3 billion active parameters + some GB for context) that can easily fit into 10 GB of ram while the not currently active experts ar
by upboundspiral 4mo ago
you can run Qwen 3.6 35B-A3B (3 billion active parameters + some GB for context) that can easily fit into 10 GB of ram while the not currently active experts are offloaded to the cpu ram with llama-cpp.