4 ms·
I see comments here about running Llama on 4090's, which is fine for local development and testing - but getting into production is a significant leap and a sig
by binarymax 3y ago
I see comments here about running Llama on 4090's, which is fine for local development and testing - but getting into production is a significant leap and a significant cost.
The thing that I keep running into in my SLA plans is concurrency. Yes, you can have a Llama 2 model running on an A100 somewhere - but that will support 1 concurrent prompt. Anything at a higher concurrency needs another GPU, or your end users will be waiting a while. Want to rent an 8 GPU machine in the cloud for inference? Be prepared to pay a lot of money for it.
- huac 3y agoyou need an inference server. I am doing ~400 tokens/sec on 7B with a 4090 with multiple concurrent (streaming!) requests. it's reasonably straightforward for me to host this and serve public requests, but would likely just be a base model -- not sure if hosting (eg) 13B chat can serve peoples' use cases
- binarymax 3y agoBut is the 7B model any good and actually production worthy for things like RAG?
- treprinum 3y agoNot as good as GPT-4 of course but nor far from 3.5 if you just need to reword whatever returned by the retrieval. It's like losing 20 IQ points which might be still better than most support interactions I had.
- huac 3y agoI'm writing a blog post with some more reasoning but my view is that it can be useful for certain simpler tasks (eg unstructured -> structured, basic summarization) and not more complex things (eg generation). The tricky thing is that finetuning makes a big difference, and while it should be possible to hotswap LoRA adapters (at some cost to performance), I haven't figured that out yet.
- treprinum 3y agoLLaMA 2 7B 8-bit can run pretty well on 64 core EPYCs which are cheaper than GPU instances. Moreover, you can periodically batch multiple users and not just run a single inference for a single user.