4 ms·
Cheapest Llama2 chatbot solution costs only $4/mon
- _flux 3y agoIt seems to suggest t3.large by default, which has 2 vCPUs and 8 GB of RAM. What kind of performance can one expect to get from such a system?
- garciasn 3y agoThey're saying it's a LLaMA-2 7B model, but I assume it's heavily quantized (Q2 or Q3 are mostly likely) because if it's got to fit in 8GB of RAM, it's going to need to be. Q2 and Q3 quantized models will require somewhere between 5 and 6GB of RAM). They might be able to go as high as a Q5 if they have extremely low overhead for the rest of the tooling (up to Q5 may fit in as little as 7GB). Performance is going to be...poor. I would expect this to be less than 1 token a second.
- ronsor 3y agoYour math doesn't quite add up. An 8-bit quantized 7B parameter model would use only ~7GB of memory. That aside, I'm able to infer a quantized 13B model on a 2-core CPU at ~2 tokens/sec, so performance would likely be fine compared to actual ChatGPT.
- smoldesu 3y agoEh, I've seen cheaper. I run Llama on a free 4-core Oracle system.
- archibaldJ 3y agowhat about 70B? What’s the lowest cost we can expect?
- scottydelta 3y agoIt's $4/month to use the image, what about the monthly cost of the t3.large server it is suggesting? The total cost comes to approx $70/month.
- win4r 3y ago[dead]
- win4r 3y ago[dead]
- olafura 3y agoThey should link to this from the amazon page: https://ai.meta.com/llama/use-policy/ https://ai.meta.com/llama/use-policy/