3 ms·
You can try Groq API for faster inference. They use custom hardware to speed up the inference. Supported open models can be found here: https://console.groq.com
by mks_shuffle 2y ago
You can try Groq API for faster inference. They use custom hardware to speed up the inference. Supported open models can be found here: https://console.groq.com/docs/models https://console.groq.com/docs/models (includes llama-70b)
- yungtriggz 2y agothanks, tried this to some mixed results. seems like they have caps on speed/rate limits etc if you havent spoken to them so might reach out