3 ms·
It currently uses only the CPU via ARM NEON intrinsics - no GPU, no ANE and no Apple Accelerate. Plan is in the future to utilize respective SIMD intrinsics fo
by ggerganov 4y ago
It currently uses only the CPU via ARM NEON intrinsics - no GPU, no ANE and no Apple Accelerate.
Plan is in the future to utilize respective SIMD intrinsics for other architectures (AVX, WASM SIMD, etc) and also add other more accurate quantization approaches. It's actually not a lot of work and I have most of the stuff ready, so hopefully soon!
Edit: AVX2 support has just been added
- lxe 4y agoAre you observing higher tokens/second throughout on Apple silicon (you’ve been mentioning 20/second) than running it in PyTorch on a CUDA GPU such as a 3090?
- ggerganov 4y agoI haven't run the original PyTorch model not a single time! I just look at the code and port it. I don't have the hardware to run it.
- simonw 4y agoAre you funded at all? If you need extra hardware I'm sure the community could make that happen.
- MacsHeadroom 4y agoA 4090 gets 30 tokens/second with LLaMA-30B, which is about 10 times faster than the 300ms/token people are reporting in these comments. (20 tokens/second on a Mac is for the smallest model, ~5x smaller than 30B and 10x smaller than 65B)
- lxe 4y agoHow are you getting 20 tokens/second? I'm getting 2.6 tokens/s on 3090 with int4 prequantized model. Is 4090 so much faster?