3 ms·
What about RTX 3080? Too little VRAM?
by hn_acc1 3mo ago
What about RTX 3080? Too little VRAM?
- roadside_picnic 3mo agoIn addition to models getting better, the quantization methods have also got much better. If you already have an RTX 3080 it's absolutely worth the time to just mess around and see how it does, experiment with different quants that fit in your VRAM. If you're purchasing I would recommend coughing up the extra cash for the 3090. If you are experimenting it's worth mentioning that the harness/tooling is very important to getting a solid experience. Herme's agent is great for running helpful agents and OpenWeb UI can get really make the experience feel on par with paid chat interfaced. A reasonable halfway step is to pay for an open model through the provider or open router. You'll get many of the benefits (especially around pricing) without needing to shell out on hardware before deciding if you like the way these models work.
- upboundspiral 3mo agoyou can run Qwen 3.6 35B-A3B (3 billion active parameters + some GB for context) that can easily fit into 10 GB of ram while the not currently active experts are offloaded to the cpu ram with llama-cpp.