5 ms·
Fast Llama 2 on CPUs with Sparse Fine-Tuning and DeepSparse
- rkwasny 3y agoVery interesting, as I understand it was fine tuned on GSM8k, so we know we can remove 60% of the model and still answer GSM8k questions Bigger question is, does the sparse model maintain any general knowledge?
- andy99 3y agoIs it compared anywhere with a smaller model fine-tuned on the same thing? Edit: I didn't see any comparisons skimming the paper. More specifically fine tuned smaller models can often be pretty good. I'd want to see how 3B and 1B etc llama models fine-tuned on the same dataset perform. Is sparsity the key here or is it just fewer parameters? The quantization part was interesting - they also quantized activations and have some adaptive quantization that accommodates outliers.
- tarruda 3y agoSeems promising. Do they say anywhere how many tokens / second is achieved on the CPU?
- littlestymaar 3y agoThey do. It's on the third graph: https://neuralmagic.com/wp-content/uploads/2023/11/CHART-Llama2-Sparse_Fine_tuned_GSM8k_Text_Generation_Performance-005-1-1536x1000.png https://neuralmagic.com/wp-content/uploads/2023/11/CHART-Lla...
- RossBencina 3y agoInteresting company. Yannic Kilcher interviewed Nir Shavit last year and they went into some depth: https://www.youtube.com/watch?v=0PAiQ1jTN5k https://www.youtube.com/watch?v=0PAiQ1jTN5k DeepSparse is on GitHub: https://github.com/neuralmagic/deepsparse https://github.com/neuralmagic/deepsparse
- anonymousDan 3y agoNir Shavit is co-author of IMO the best book on concurrent programming: https://dl.acm.org/doi/book/10.5555/2385452 https://dl.acm.org/doi/book/10.5555/2385452
- mark_l_watson 3y agoThanks for the GitHub link, I want to try this in a M2 with 32G memory. It will “waste” the GPU and neural cores, but might still run fast. Off topic: I so much appreciate everyone’s work Getting LLMs running on inexpensive hardware! I have been having crazy amounts of fun with Ollama and a wide variety of tuned models. And, so fast!
- arkmm 3y agoI might be missing something but it seems like most of the speed up is from quantization which is commonly used already, and the CPU instance used here isn't that much cheaper (~10-15%?) than a GPU instance that could run the model. For high utilization workloads the extra throughput might be useful though.
- yunohn 3y agoI think you are missing something, or I am. If we observe the performance comparison graph (1), 2.8->9 tok/s is achieved via quantization, but the remaining jumps from 9->16.6->24.6 tok/s are achieved from the sparse fine-tuning. (1) https://neuralmagic.com/wp-content/uploads/2023/11/CHART-Llama2-Sparse_Fine_tuned_GSM8k_Text_Generation_Performance-005-1-1536x1000.png https://neuralmagic.com/wp-content/uploads/2023/11/CHART-Lla...
- jonatron 3y agoHetzner do cheap servers, but don't have many GPUs. If you're looking at saving money by running on CPU, you shouldn't be looking at one of the most expensive server providers.
- thelastparadise 3y agoHow is the $ per inference, say 4k tokens, on a Hetzner box vs an A100? Not all tasks require low latency.
- jonatron 3y agoI'd like to know the answer to that too. I shouldn't have implied that CPU inference on Hetzner is cheaper than GPU on AWS, when I don't have any idea on the cost of either.
- mikeravkine 3y agoHetzner offers incredibly cheap ARM machines in the Falkenstein DC, for 25Eur a month you can snag the top of the line with 16 vCPU and 32GB RAM. If your usecase fits inside that 32GB (no 70B models, sadly) the price to performance of a GGUF Q4KM is really attractive on this setup.
- leobg 3y agoAny way to run a fine tuned Mistral model?
- jakey_bakey 3y agoDoes anyone have a good resource on fine-tuning using open-source LLMs?
- shultays 3y agoWhat "no drop in accuracy" means? Is it like they do some kind of lossless compression and it is guaranteed to behave same? Or are they claiming that as a result of a (subjective?) test that measures accuracy? If so how does such test works?
- halflings 3y agoThe page fully explains what they mean by this, showing results on benchmarks etc.
- thepangolino 3y ago[dead]
- avipars 3y agoIs it RAM-usage heavy? I tried other models and they crashed because it stored too much data in RAM
- wills_forward 3y agoThis seems like a big step forward in terms of running specific use-case trained inference locally, right? At least given current hardware generally deployed in business.
- lawlessone 3y agoIs this the same as the Fast feed forwards post from yesterday? If not, could both be applied?
- visarga 3y agoThey are alternative approaches.
- buildbot 3y agoI did a but of digging but was unable to find what kind if sparsity this applies - unstructured? Semi-structured? Block sparsity?
- jillbore 3y agoThis is about seeing inference benefits from unstructured sparsity.