9 ms·
Microsoft researchers developed a hyper-efficient AI model that can run on CPUs
- hu3 1y agoRepo with demo video and benchmark: https://github.com/microsoft/BitNet https://github.com/microsoft/BitNet "...It matches the full-precision (i.e., FP16 or BF16) Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption..." https://arxiv.org/abs/2402.17764 https://arxiv.org/abs/2402.17764
- Animats 1y agoThat essay on the water cycle makes no sense. Some sentences are repeated three times. The conclusion about the water cycle and energy appears wrong. And what paper is "Jenkins (2010)"? Am I missing something, or is this regressing to GPT-1 level?
- yorwba 1y agoThey should probably redo the demo with their latest model. I tried the same prompt on https://bitnet-demo.azurewebsites.net/ https://bitnet-demo.azurewebsites.net/ and it looked significantly more coherent. At least it didn't get stuck in a loop.
- int_19h 1y ago2B parameters should be in the ballpark of GPT-2, no?
- godelski 1y ago> "...It matches the full-precision (i.e., FP16 or BF16) Wait... WHAT?! When did //HALF PRECISION// become //FULL PRECISION//? FWIW, I cannot find where you're quoting from. I cannot find "matches" on TFA nor the GitHub link. And in the paper I see 3.2 Inference Accuracy The bitnet.cpp framework enables lossless inference for ternary BitNet b1.58 LLMs. To evaluate inference accuracy, we randomly selected 1,000 prompts from WildChat [ ZRH+24 ] and compared the outputs generated by bitnet.cpp and llama.cpp to those produced by an FP32 kernel. The evaluation was conducted on a token-by-token basis, with a maximum of 100 tokens per model output, considering an inference sample lossless only if it exactly matched the full-precision output.
- falcor84 1y agoWhy do they call it "1-bit" if it uses ternary {-1, 0, 1}? Am I missing something?
- sambeau 1y agoMaybe they are rounding down from 1.5-bit :)
- BuyMyBitcoins 1y agoClassic Microsoft naming shenanigans.
- 1970-01-01 1y agoIt's not too late to claim 1bitdotnet.net before they do.
- DecentShoes 1y agoLLM Series One S and X
- Maxious 1y agohttps://compilade.net/blog/ternary-packing https://compilade.net/blog/ternary-packing is a good explainer (previous discussion https://news.ycombinator.com/item?id=42329307 https://news.ycombinator.com/item?id=42329307)
- falcor84 1y agoThanks, but I've skimmed through both and couldn't find an answer on why they call it "1-bit".
- AzN1337c0d3r 1y agoThe original BitNet paper (https://arxiv.org/pdf/2310.11453 https://arxiv.org/pdf/2310.11453) BitNet: Scaling 1-bit Transformers for Large Language Models was actually binary (weights of -1 or 1), but then in the follow-up paper they started using 1.58bit weights (https://arxiv.org/pdf/2402.17764 https://arxiv.org/pdf/2402.17764) The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits This seems to be first source of the confounding of "1-bit LLM" and ternary weights that I could find. In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in which every single parameter (or weight) of the LLM is ternary {-1, 0, 1}.
- justanotheratom 1y agoSuper cool. Imagine specialized hardware for running these.
- LargoLasskhyfv 1y agoIt already exists. Dynamically reconfigurable. Some smartass designed it alone on ridiculously EOL'd FPGAs. Meanwhile ASICs in small batches without FPGA baggage were produced. Unfortunately said smartass is under heavy NDA. Or luckily, because said NDA paid very well for him.
- djmips 1y agoNicely done!
- LargoLasskhyfv 1y agoWas actually sort of a sideways pivot, and hard for me to do, because of the involved mathematics. Initially it was more of general 'architecture astronautics' in the context of dynamically reconfigurability/systolic arrays/transport triggered architecture/VLIW, which got me some nice results. Having read and thought much about balanced ternary hardware, and 'playing' with that, while also reading how this could be favourably applicable to ML lead to that 'pivot'. A few years before 'this', I might add. Now I can 'play' much more relaxed and carefree, to see what else I can get out of this :-)
- llama_drama 1y agoI wonder if instructions like VPTERNLOGQ would help speed these up
- ilrwbwrkhv 1y agoThis will happen more and more. This is why NVidia is rushing to get CUDA a software level lock-in otherwise their stock will go the way of Zoom.
- soup10 1y agoi agree, no matter how much wishful thinking jensen sells to investors about paradigm shifts the days of everyone rushing out to get 6 figure tensor core clusters for their data center probably won't last forever.
- bigyabai 1y agoIf Nvidia was at all in a hurry to lock-out third-parties, then I don't think they would support OpenCL and Vulkan compute, or allow customers to write PTX compilers that interface with Nvidia hardware. In reality, the demand for highly parallelized compute simply blindsided OEMs. AMD, Intel and Apple were all laser-focused on raster efficiency, none of them have a GPU architecture optimized for GPGPU workloads. AMD and Intel don't have competitive fab access and Apple can't sell datacenter hardware to save their life; Nvidia's monopoly on attractive TSMC hardware isn't going anywhere.
- mlinhares 1y agoThe profit margins on Macs must be insane because it just doesn’t make sense at all Apple just doesn’t give a fuck about data center workloads when they have some of the best ARM CPUs and whole packages on the market.
- bigyabai 1y agoIf Xserve is any basis of comparison, Apple struggles to sell datacenter hardware in the best of markets. The competition is too hot nowadays, and Apple likely knows the investment wouldn't be worth it. ARM CPUs are available from Ampere and Nvidia now, Apple Silicon would have to differentiate itself more than it does on mobile. After a certain point, it probably does come down to the size of the margins on consumer hardware.
- 1970-01-01 1y ago..and eventually the Skynet Funding Bill was passed.
- stogot 1y agoThe pricing war will continue to rock bottom
- Jedd 1y agoI think almost all the free LLMs (not AI) that you find on hf can 'run on CPUs'. The claim here seems to be that it runs usefully fast on CPU. We're not sure how accurate this claim is, because we don't know how fast this model runs on a GPU, because: > Absent from the list of supported chips are GPUs [...] And TFA doesn't really quantify anything, just offers: > Perhaps more impressively, BitNet b1.58 2B4T is speedier than other models of its size — in some cases, twice the speed — while using a fraction of the memory. The model they link to is just over 1GB in size, and there's plenty of existing 1-2GB models that are quite serviceable on even a mildly-modern CPU-only rig.
- deleted 1y ago[deleted]
- sheepscreek 1y agoIf you click the demo link, you can type a live prompt and see it run on CPU or GPU (A100). From my test, the CPU was laughably slower. To my eyes, it seems comparable to the models I can run with llama.cpp today. Perhaps I am completely missing the point of this.
- esafak 1y agoIs there a library to distill bigger models into BitNet?
- timschmidt 1y agoI could be wrong, but my understanding is that bitnet models have to be trained that way.
- babelfish 1y agoThey don't have to be trained that way! The training data for 1-bit LLMs is the same as for any other LLM. A common way to generate this data is called 'model distillation', where you take completions from a teacher model and use them to train the child model (what you're describing)!
- timschmidt 1y agoMaybe I wasn't clear, I think you've misunderstood me. I understand that all sorts of LLMs can be trained using a common corpus of data. But my understanding is that the choice of creating a bitnet LLM must be made at training time, as modifications to the training algorithms are required. In other words, an existing FP16 model cannot be quantized to bitnet.
- babelfish 1y agoAh yes, definitely misunderstood you, my bad
- ein0p 1y agoThis is over a year old. The sky did not come down, everyone didn't switch to this in spite of the "advantages". If you look into why, you'll see that it does, in fact, affect the metrics, and some more than others, and there is no silver bullet.
- justanotheratom 1y agoare you predicting, or is there already a documented finding somewhere?
- ein0p 1y agoTake a look at their own paper or at many attempts to train something large with this. There's no replacement for displacement. If this actually worked without quality degradation literally everyone would be using this.
- yorwba 1y agoThe 2B4T model was literally released yesterday, and it's both smaller and better than what they had a year ago. Presumably the next step is that they get more funding for a larger model trained on even more data to see whether performance keeps improving. Of course the extreme quantization is always going to impact scores a bit, but if it lets you run models that otherwise wouldn't even fit into RAM, it's still worth it.
- imtringued 1y agoAQLM, EfficientQAT and ParetoQ get reasonable benchmark scores at 2-bit quantization. At least 90% of the original unquantized scores.
- zamadatix 1y ago"Parameter count" is the "GHz" of AI models: the number you're most likely to see but least likely to need. All of the models compared (in the table on the huggingface link) are 1-2 billion parameters but the models range in actual size by more than a factor of 10.
- int_19h 1y agoBecause of different quantization. However, parameter count is generally the more interesting number so long as quantization isn't too extreme (as it is here). E.g. FP32 is 4x the size of 8-bit quant, but the difference is close to non-existent in most cases.
- orbital-decay 1y ago>so long as quantization isn't too extreme (as it is here) This is true for post-training quantization, not for quantization-aware training, and not for something like BitNet. Here they claim comparable performance per parameter count as normal models, that's the entire point.
- charcircuit 1y agoTPS is the Ghz of AI models. Both are related to the the propagation time of data.
- idonotknowwhy 1y agoThen i guess vocab is the IPC. 10k mistral tokens are about 8k llama3 tokens
- nodesocket 1y agoThere are projects working on distributed LLMs, such as exo[1]. If they can crack the distributed problem fully and get performance it’s a game changer. Instead of spending insane amounts on Nvidia GPUs, can just deploy commodity clusters of AMD EPYC servers with tons of memory, NVMe disks, and 40G or 100G networking which is vastly less expensive. Goodbye Nvidia AI moat. [1] https://github.com/exo-explore/exo https://github.com/exo-explore/exo
- lioeters 1y agoDo you think this is inevitable? It sounds like, if distributed LLMs are technically feasible to achieve, it will eventually happen. Maybe that's an unknown whether it can be solved at all, but I imagine there are enough people working on the problem that they will find a break-through one way or the other. LLMs themselves could participate in solving it. Edit: Oh I just saw the Git repo: > exo: Run your own AI cluster at home with everyday devices. So the "distributed problem" is in the process of being solved. Impressive.
- instagraham 1y ago> it’s openly available under an MIT license and can run on CPUs, including Apple’s M2. Weird comparison? The M2 already runs 7 or 13gb LLama and Mistral models with relative ease. The M-series and Macbooks are so ubiquitous that perhaps we're forgetting how weak the average CPU (think i3 or i5) can be.
- nine_k 1y agoThe M-series have a built-in GPU and unified RAM accessible to both. Running a model on an M-series chip without using the GPU is, imho, pointless. (That said, it's still a long shot from an H100 with a ton of VRAM, or from a Google TPU.) If a model can be "run on a CPU", it should acceptably run on a general-purpose 8-core CPU, like an i7, or i9, or a Ryzen 7, or even an ARM design like a Snapdragon.