3 ms·
why? it's mostly reads. the weights are static.
by RobMurray 6mo ago
why? it's mostly reads. the weights are static.
- bigyabai 6mo agollama-cpp's process is, but macOS itself will swap hard when 10-14gb of memory is paged for LLM inference. Dense models especially would thrash zram.