3 ms·
https://compilade.net/blog/ternary-packing https://compilade.net/blog/ternary-packing is a good explainer (previous discussion https://news.ycombinator.com/item
by Maxious 1y ago
https://compilade.net/blog/ternary-packing https://compilade.net/blog/ternary-packing is a good explainer (previous discussion https://news.ycombinator.com/item?id=42329307 https://news.ycombinator.com/item?id=42329307)
- falcor84 1y agoThanks, but I've skimmed through both and couldn't find an answer on why they call it "1-bit".
- AzN1337c0d3r 1y agoThe original BitNet paper (https://arxiv.org/pdf/2310.11453 https://arxiv.org/pdf/2310.11453) BitNet: Scaling 1-bit Transformers for Large Language Models was actually binary (weights of -1 or 1), but then in the follow-up paper they started using 1.58bit weights (https://arxiv.org/pdf/2402.17764 https://arxiv.org/pdf/2402.17764) The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits This seems to be first source of the confounding of "1-bit LLM" and ternary weights that I could find. In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in which every single parameter (or weight) of the LLM is ternary {-1, 0, 1}.
- LeonB 1y agoIt’s “1-bit, for particularly large values of ‘bit’”
- biomcgary 1y agoShould be 1-trit.
- taneq 1y agoThat’s pretty cool. :) One thing I don’t get is why do multiple operations when a 243-entry lookup table would be simpler and hopefully faster?
- compilade 1y agoBecause lookup tables are not necessarily faster compared to 8-bit SIMD operations, at least when implemented naïvely. Lookup tables can be fast, but it's not simpler, see T-MAC https://arxiv.org/abs/2407.00088 https://arxiv.org/abs/2407.00088 (Note that all comparisons with `llama.cpp` were made before I introduced the types from https://github.com/ggml-org/llama.cpp/pull/8151 https://github.com/ggml-org/llama.cpp/pull/8151 where the 1.6-bit type uses the techniques described in the aforementioned blog post). I wanted to try without lookup tables to at least have a baseline, and also because the fixed point packing idea lent itself naturally to using multiplications by powers of 3 when unpacking.
- taneq 1y agoThanks for taking the time to reply! I haven’t done any serious low level optimisation on modern CPUs so most of my intuitions are probably way out of date.