4 ms·
But this doesn’t apply to self-hosted, o?
by ComputerGuru 3y ago
But this doesn’t apply to self-hosted, o?
- lyjackal 3y agoIt does. LLMs are most efficient when running large batches, so the gpu cost is super high if you’re underutilizing it. It will cost more than a cloud provider like open ai who has the volume to keep their GPUs saturated
- jay-barronville 3y agoYup. It’s also important to mention that OpenAI enjoys the luxury of having large clusters of H100s (the last time I checked).