3 ms·
They're saying this method essential does not, even when mixed with low rank models on top. "Notably, while the original BF16 model requires per-layer CPU offlo
by timnetworks 2y ago
They're saying this method essential does not, even when mixed with low rank models on top. "Notably, while the original BF16 model
requires per-layer CPU offloading on the 16GB laptop 4090, our INT4 model fits entirely in GPU memory, resulting in a 10.1× speedup by avoiding offloading."
This is the whole magic, the rest of the workflow doesn't need to unload and flush memory, causing big delays for jobs.