3 ms·
I've only read the abstract but they don't mention quantizing the weights or otherwise trying to shrink the model in any way. They're claiming to be able to ef
by zackangelo 2y ago
I've only read the abstract but they don't mention quantizing the weights or otherwise trying to shrink the model in any way.
They're claiming to be able to efficiently run larger models without loading the entire thing into GPU memory. If they're using the same weights, the same architecture and just using tensor parallel operations to perform the forward pass that would imply no loss in quality.
I'm sure there are trade-offs but they're not clear by just looking at the abstract.
- tgtweak 2y agoI read it like this too - no drop in weights or model quality just optimizing the lower boundaries of performance when you are splitting from vram to ram to disk (or network).