3 ms·
You can get llama CPP or kobold.cpp binaries and load a quantized model right into them on the CPU only, no need to install Python or have an Nvidia GPU.
by dwringer 3y ago
You can get llama CPP or kobold.cpp binaries and load a quantized model right into them on the CPU only, no need to install Python or have an Nvidia GPU.