3 ms·
I might be missing something but it seems like most of the speed up is from quantization which is commonly used already, and the CPU instance used here isn't th
by arkmm 3y ago
I might be missing something but it seems like most of the speed up is from quantization which is commonly used already, and the CPU instance used here isn't that much cheaper (~10-15%?) than a GPU instance that could run the model. For high utilization workloads the extra throughput might be useful though.
- yunohn 3y agoI think you are missing something, or I am. If we observe the performance comparison graph (1), 2.8->9 tok/s is achieved via quantization, but the remaining jumps from 9->16.6->24.6 tok/s are achieved from the sparse fine-tuning. (1) https://neuralmagic.com/wp-content/uploads/2023/11/CHART-Llama2-Sparse_Fine_tuned_GSM8k_Text_Generation_Performance-005-1-1536x1000.png https://neuralmagic.com/wp-content/uploads/2023/11/CHART-Lla...
- jonatron 3y agoHetzner do cheap servers, but don't have many GPUs. If you're looking at saving money by running on CPU, you shouldn't be looking at one of the most expensive server providers.
- thelastparadise 3y agoHow is the $ per inference, say 4k tokens, on a Hetzner box vs an A100? Not all tasks require low latency.
- jonatron 3y agoI'd like to know the answer to that too. I shouldn't have implied that CPU inference on Hetzner is cheaper than GPU on AWS, when I don't have any idea on the cost of either.
- mikeravkine 3y agoHetzner offers incredibly cheap ARM machines in the Falkenstein DC, for 25Eur a month you can snag the top of the line with 16 vCPU and 32GB RAM. If your usecase fits inside that 32GB (no 70B models, sadly) the price to performance of a GGUF Q4KM is really attractive on this setup.
- londons_explore 3y agoWith two/three instances, you can probably fit a 70B model into RAM, and you don't need super low latency between models to be able to do inference split layerwise between machines.
- fbdab103 3y agoAre there instructions for this distributed inference somewhere? Can I do this out of the box with llamacpp or similar?
- londons_explore 3y agoDon't think so. I suspect it would require quite in-depth surgery of llamacpp to add in the ability to send activations over the internet and pipeline stuff to keep all the cores busy.