3 ms·
It didn't take long for perplexity, anyscale, together.ai, groq, deepinfra, or lepton to all host mistral's 8x7B model, both faster and cheaper then Mistral's o
by declaredapple 3y ago
It didn't take long for perplexity, anyscale, together.ai, groq, deepinfra, or lepton to all host mistral's 8x7B model, both faster and cheaper then Mistral's own api.
https://artificialanalysis.ai/models/mixtral-8x7b-instruct/hosts https://artificialanalysis.ai/models/mixtral-8x7b-instruct/h...
- sillysaurusx 3y agoHosting a 7B model is completely different than hosting a 150B+ model. I thought this would be obvious, but I should have been explicit.
- declaredapple 3y agoIt's not really. And 8x7B is not a 7B model, it's a MoE that's closer to 60B that has to be kept in memory, and uses 2 experts per token so it runs at 15B speeds. All of the current frameworks support MoE and sharding among GPUs so I don't see what the issue is.