3 ms·
You got me interested, a rough table of model size to instance type to spot $/hr would help a lot. The 0.5B example is CPU only, so it doesn't say much about wh
by dsemakin 24d ago
You got me interested, a rough table of model size to instance type to spot $/hr would help a lot. The 0.5B example is CPU only, so it doesn't say much about what a 7B or 70B actually costs.
- deleted 24d ago[deleted]
- paguasmar 24d agoYou're right. A CPU only model is definitely is not a real prod workload where you'd need GPU machines. Here is a rough table of model size to instance type to spot $/hr: For up to 8B, you can use a c6i.xlearge which costs as spot about $0.4/hr. For up to 70B, you can use a g5.12xlearge that costs you about $2/hr. Besides the GPU machine be aware that you need a controller instance to route the requests and scale. That has a fixed cost of up to $0.08/hr. When not in use at all, just type veloxml down --all and it tears down the controller too for true $0/hr. We're currently running larger LLM models benchmarks to add to the README this week.