4 ms·
Slightly off topic, here is the best local llama.cpp wrapper I've run into: https://github.com/Mozilla-Ocho/llamafile https://github.com/Mozilla-Ocho/llamafile
by discardedrefuse 3y ago
Slightly off topic, here is the best local llama.cpp wrapper I've run into:
https://github.com/Mozilla-Ocho/llamafile https://github.com/Mozilla-Ocho/llamafile
You can download any .gguf model (not just the ones in their examples) and run it locally (as long as you have the ram for it). I was running 7B models with ease on an old FX8350 and now 13B models on a 5600X (32GB RAM on both machines).
This wrapper spins up a local web server that runs a simple web frontend to use immediately with no code, but also exposes an OpenAI compatible API for dev work and alt frontends (like SillyTavern).
- 3abiton 3y agoWhat's the speed? I heard of llamafile before.
- discardedrefuse 3y agoI wasn't really timing it, but it felt fine, especially since it prints out while generating. The 13B model on the 5600X is about 10 seconds. Tho once the conversation gets too long (about 10 replies) it will go sideways. I gotta play with the default settings. IIRC the 7B on the Fx8350 was closer to 25 seconds per response.