3 ms·
Breaking the 1.58-bit Barrier for Ternary LLMs
- Kevcmk 9d agoWoah. Good science.
- kadushka 9d ago[flagged]
- NooneAtAll3 9d agoThis is the only time "1.58 bit" phrase makes more sense than "1 trit" Who knew that if you actually look at information entropy you can pack stuff better!
- plqbfbv 9d agoVery interesting, I was just exploring this to hopefully fit one of the latest quantized models in 16GB of VRAM.
- om8 9d agoTernary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
- om8 9d agoIf you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
- janalsncm 9d agoPTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU. If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.
- mitxela 9d agowhich is important though since sending it across the wire over and over and over is actually the main bottleneck.
- Kerbonut 9d agoWire typically means internet connection, and it’s hardly the bottleneck
- deleted 9d ago[deleted]
- 317070 9d agoin the case of large language models, the wire is the communication of your parameters between your layers of memory that is often the bottleneck. To do a forward pass, you need to use all parameters once, and so the communication between the compute and the storage is the bottleneck, and that bottleneck is also a bunch of wires.
- mitxela 9d agoThe other bottleneck is the amount of fast storage, which compression also improves.
- om8 9d ago> If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire. That’s why you need to use efficient gemm kernels like FLUTE for inference. They are ~as good as what you can do with ternary quantization.
- WithinReason 9d agoAnd storing it in memory. Memory is expensive.
- kadushka 9d ago[dead]
- infogulch 9d agoSo they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat. If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
- kadushka 9d agoBy “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
- montroser 9d agoWell, you could train directly at this bitrate.
- brookst 9d agoNot an expert, but doesn’t that produce lower quality results, the same way a 1MP image isn’t lower quality than a 20mp image downscaled to 1MP? (Everything else equal)
- Vetch 9d agoQAT, which bitnet training is a form of, helps a ton in preserving accuracy at such low bits per parameter. There are also better quantization approaches that try to preserve the most sensitive weights† but are computationally expensive and so not typically done. Another complementary option is, if the model is fast enough, we should be able to push up correctness by self-consistency voting at close to T=1. Smart/fast Zero-shot classifiers like the recent Jev could help with aggregation across answers too, extending applicability. †Every paper I've read estimates the average information content of transformer LLMs at about 3-4 bits per parameter. Curiously, biological synapses are also estimated to be about 4-5 bits per synapse, possibly a bit lower.
- kadushka 9d ago
- wgd 9d agoOnly a presence bitmap? If we're contemplating packing schemes I'm tempted to write a paper that uses arithmetic coding to squeeze out a few more centi-bits.
- akoboldfrying 9d agoAgreed. One possible objection might be that they need fast random access to weights, but I skimmed parts of the paper and it looks like they process 128 entries at a time, which to me sounds like it should be amenable to better compression: short enough that better compression results could still be efficiently cached in faster local RAM, long enough that better compression would save useful amounts of memory per block.
- yalok 9d agosounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference. And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done... 0. https://arxiv.org/pdf/2402.17764 https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
- Marchant_hq 9d agoPushing past log2(3) for real. This could drastically shrink LLMs for embedded systems, making them truly portable.
- kittikitti 9d agoThank you for sharing this. I like to test out running LLM's on edge computing with limited RAM and GPU/CPU so this research will have practical implications on my activities. I also appreciated how the authors formulated 1.58 (it's log_2(3)) because that was embarrassingly confusing for me when I was first introduced to ternary LLM's.
- explainit2me 9d agoSo this compression is only pertinent to the LLM file format? In memory it'd have to be expanded into the 1.58-bit form - 5 trits per byte.
- pieter3d 9d agoIt also means you can read them faster, more parameters per second during an inference which tends to be memory bandwidth limited on most systems. Thus faster inferences
- c7b 9d ago> We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout I honestly assumed that's how they already work. I have to admit that I even explained it like that to a friend. Why on earth wouldn't you design it like that from the start (talking about the adaptive, not the measure part; just sacrifice a few bits to clarify your encoding and save a ton of bits)?
- ant6n 9d agoternary is totally losslessly compressed anyway. Why not just use an 8-bit LUT to encode the 256 most common ternary vectors with 6 components. That means of the possible 729 possible such vectors, you can only represent 256 different ones. You have to do more aggressive rounding, but at least the scheme is very simple to decompress and stream.
- akoboldfrying 9d agoI think the weights are iid distributed, so all 729 patterns will be roughly equally likely. That doesn't make this a bad idea though -- it just means there's no point trying to select the most common 256 to keep, since any 256 will be roughly as good.
- ant6n 8d agoIn the article it said that in ternary, the majority of weights are 0. The components may be independent, but any group of weights won't be evenly distributed across all probable occurences.
- akoboldfrying 7d agoGood point, I was wrong. Groups of weights having more zeros will be more likely, so should be favoured. Due to independence there won't be any meaningful difference in frequencies between two groups of weights that have the same number of zeros, but that doesn't invalidate the above.
- CodesInChaos 9d agoI'm surprised that a variable length encoding like this is usable directly as in memory format and not just as storage/transfer format.
- locitra 9d ago[flagged]
- itsmeduncan 9d ago[flagged]
- quillan_audits 9d ago[flagged]