3 ms·
The really big thing is the model weights, which are a read workload not a write one, it won't affect an SSD's write endurance. The write workloads are just th
by mappu 2mo ago
The really big thing is the model weights, which are a read workload not a write one, it won't affect an SSD's write endurance.
The write workloads are just the context and any K/V cache - llama.cpp does not mmap those to disk, so they would remain in memory or VRAM as space affords.
- walrus01 2mo agoI plan to give it a try in a day or two with llama-server from the main branch compiled today, when my Q8 GGUF download of K3 finishes, on a system with 256GB (should be more than ample for context and KV cache and a moderate chunk of the whole 1.6TB). If it works it's going to be sloooooooow as hell, but it'll be an interesting data point to see just how slow.