3 ms·
easiest is probably with ollama [0]. I think the ollama API is OpenAI compatible. [0]https://ollama.com/ https://ollama.com/
by ru552 2y ago
easiest is probably with ollama [0]. I think the ollama API is OpenAI compatible.
[0]https://ollama.com/ https://ollama.com/
- talldayo 2y agoMost inference servers are OpenAI-compatibile. Even the "official" llama-cpp server should work fine: https://github.com/ggerganov/llama.cpp/blob/master/examples/server/README.md https://github.com/ggerganov/llama.cpp/blob/master/examples/...
- pants2 2y agoOllama runs locally. What's the best option for calling the new Mixtral model on someone else's server programmatically?
- Arcuru 2y agoOpenrouter lists several options: https://openrouter.ai/models/mistralai/mixtral-8x22b https://openrouter.ai/models/mistralai/mixtral-8x22b