4 ms·
How are you currently hosting your LLaMA 2? Any tips, tricks or advice?
by armini 3y ago
How are you currently hosting your LLaMA 2? Any tips, tricks or advice?
- syntaxing 3y agoIt depends on your needs. For instance, do you want to host an API or do you want to have a front end like chatGPT? Chances are, text-generation-webui [1] should get you pretty close to hosting it yourself. You simply clone the repo, download the model from huggingface using the included helper (download-model.py) and fire up the server with server.py. You can connect to it by SSH port tunneling on port 7860 (there's other way like Ngrok but SSH tunneling is the easiest and secure). As for hosting, I found that runpod [2] has been the cheapest (not affiliated, just a user). All the other services tend to add up more than them when you include bandwidth and storage. There's some tutorials online [3] but a lot of them use the quantized version. You should be able to fit the original 70B with "load_in_8bit" on one A100 80GB. [1] https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui [2] https://www.runpod.io/ https://www.runpod.io/ [3] https://gpus.llm-utils.org/running-llama-2-on-runpod-with-oobaboogas-text-generation-webui/ https://gpus.llm-utils.org/running-llama-2-on-runpod-with-oo...
- robertnishihara 3y agoIf you want to query the Llama-2 models, you can use Anyscale Endpoints [1]. Note: I work on this :) Llama-2-70B is $1 / million tokens, which is the most cost-efficient on the market that I'm aware of. [1] https://app.endpoints.anyscale.com/ https://app.endpoints.anyscale.com/
- TeMPOraL 3y agoTried to plug it in to my favorite chat frontend (TypingMind), bounced off CORS. Is this something you can do something about?
- andrewmunn 3y agoHow do you keep the cost down?
- zo1 3y agoCan we supply our own fine-tuned models? Edit. I'm sure it's answered on your site but sometimes it's better to include it right here! :)