4 ms·
hmm, I'm getting the same results - but I see on M1 with a 7b model we should expect ~10x faster prompt processing https://github.com/ggml-org/llama.cpp/discus
by brrrrrm 1y ago
hmm, I'm getting the same results - but I see on M1 with a 7b model we should expect ~10x faster prompt processing
https://github.com/ggml-org/llama.cpp/discussions/4167 https://github.com/ggml-org/llama.cpp/discussions/4167
I wonder if it's the encoder that isn't optimized?