4 ms·
Kinda bummed, I get why he used Ollama but I feel like using llama cpp directly would provide better and more consistent results
by syntaxing 1y ago
Kinda bummed, I get why he used Ollama but I feel like using llama cpp directly would provide better and more consistent results
- mkl 1y agoAs the article describes, most of this was done with llama.cpp, not Ollama.
- syntaxing 1y agoAhh good catch, I didn’t notice if you scroll lower, he has the llama cpp results. The ollama-benchmark repo name is a misnomer.
- geerlingguy 1y agoI'm slowly migrating all my testing to https://github.com/geerlingguy/beowulf-ai-cluster https://github.com/geerlingguy/beowulf-ai-cluster
- RossBencina 1y agoI heard that ik_llama.cpp performs better for CPU use: https://github.com/ikawrakow/ik_llama.cpp/ https://github.com/ikawrakow/ik_llama.cpp/