4 ms·
Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
by om8 16d ago
Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
- om8 16d agoIf you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
- janalsncm 16d agoPTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU. If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.
- mitxela 16d agowhich is important though since sending it across the wire over and over and over is actually the main bottleneck.
- Kerbonut 15d agoWire typically means internet connection, and it’s hardly the bottleneck
- deleted 15d ago[deleted]
- 317070 15d agoin the case of large language models, the wire is the communication of your parameters between your layers of memory that is often the bottleneck. To do a forward pass, you need to use all parameters once, and so the communication between the compute and the storage is the bottleneck, and that bottleneck is also a bunch of wires.
- mitxela 15d agoThe other bottleneck is the amount of fast storage, which compression also improves.
- om8 15d ago> If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire. That’s why you need to use efficient gemm kernels like FLUTE for inference. They are ~as good as what you can do with ternary quantization.
- WithinReason 15d agoAnd storing it in memory. Memory is expensive.