3 ms·
Oobagooda and other front ends and similar projects have in my testing had upwards of a 50% difference in inference speed on the same model and settings, So ben
by UnlockedSecrets 3y ago
Oobagooda and other front ends and similar projects have in my testing had upwards of a 50% difference in inference speed on the same model and settings, So benchmarks are still useful.
- brucethemoose2 3y agoOoba is an outlier, and has tons of overhead over llama.cpp and llama-cpp-python for some reason. Most llama.cpp openai servers are pretty close to vanilla llama.cpp, albeit without the batching support.