3 ms·
Tokens are expensive, sure, but I don't even _want_ Ollama to run inference for me. Ollama gives me, essentially, a wrapper for llama.cpp and convenient hostin
by hephaes7us 1y ago
Tokens are expensive, sure, but I don't even _want_ Ollama to run inference for me.
Ollama gives me, essentially, a wrapper for llama.cpp and convenient hosting where I can download models.
I'm happy to pay for the bandwidth, plus a premium to cover their running this service.
I'm furthermore happy to pay a small charge to cover the development that they've done and continue to do to make local-inference easy for me.