3 ms·
I've been seeing a lot of projects that let one use large models on machines with small amounts of memory. They seem to all be doing significant quantization an
by hatthew 2mo ago
I've been seeing a lot of projects that let one use large models on machines with small amounts of memory. They seem to all be doing significant quantization and/or expert streaming. What's the benefit of these projects over something like taking an unsloth quant and running llama.cpp with appropriate flags (-cmoe/-mmap) to manage VRAM vs RAM vs SSD?