3 ms·
Prior to LLaMA 2, I would have agreed with you but LLaMA 2 is a game changer. The 70B performance is probably between 3.5 and 4. But running it personally isn't
by syntaxing 3y ago
Prior to LLaMA 2, I would have agreed with you but LLaMA 2 is a game changer. The 70B performance is probably between 3.5 and 4. But running it personally isn't cheap. The cheapest I found is about $4/hr to run the whole thing. I only spend around $3 on average a month on GPT-3.5 API for my personal stuff.
- armini 3y agoHow are you currently hosting your LLaMA 2? Any tips, tricks or advice?
- syntaxing 3y agoIt depends on your needs. For instance, do you want to host an API or do you want to have a front end like chatGPT? Chances are, text-generation-webui [1] should get you pretty close to hosting it yourself. You simply clone the repo, download the model from huggingface using the included helper (download-model.py) and fire up the server with server.py. You can connect to it by SSH port tunneling on port 7860 (there's other way like Ngrok but SSH tunneling is the easiest and secure). As for hosting, I found that runpod [2] has been the cheapest (not affiliated, just a user). All the other services tend to add up more than them when you include bandwidth and storage. There's some tutorials online [3] but a lot of them use the quantized version. You should be able to fit the original 70B with "load_in_8bit" on one A100 80GB. [1] https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui [2] https://www.runpod.io/ https://www.runpod.io/ [3] https://gpus.llm-utils.org/running-llama-2-on-runpod-with-oobaboogas-text-generation-webui/ https://gpus.llm-utils.org/running-llama-2-on-runpod-with-oo...
- robertnishihara 3y agoIf you want to query the Llama-2 models, you can use Anyscale Endpoints [1]. Note: I work on this :) Llama-2-70B is $1 / million tokens, which is the most cost-efficient on the market that I'm aware of. [1] https://app.endpoints.anyscale.com/ https://app.endpoints.anyscale.com/
- TeMPOraL 3y agoTried to plug it in to my favorite chat frontend (TypingMind), bounced off CORS. Is this something you can do something about?
- andrewmunn 3y agoHow do you keep the cost down?
- zo1 3y agoCan we supply our own fine-tuned models? Edit. I'm sure it's answered on your site but sometimes it's better to include it right here! :)
- jimmcslim 3y agoOut of curiosity and if you are happy to share, what is your 'personal stuff'?
- SOLAR_FIELDS 3y agoAs a counter reference, for my work I use it to code (for-4) and it has been between $70 and $200 per month depending on how heavily I use it
- syntaxing 3y agoGPT-4 is significantly more expensive so I can definitely see you spending that amount. For really complex stuff, I switch over the GPT-4 and it will cost me almost $3 a "question" (as in going from the beginning to solving it). Honestly worth it since it solves my problem but it adds up quick so I try to stick with 3.5 when I can.
- blorenz 3y agoCan’t you get by with ChatGPT-4 for these personal assistant type questions? That’s what I do and my 20 a month goes a long way. I’d be interested to see if I am missing out on anything using GPT to is way in contrast to the API.
- facia 3y ago[dead]
- SOLAR_FIELDS 3y agoI use it with a tool that is wired into my terminal that changes my files for me [1]. That alone makes me several times more productive compared to copy pasting back and forth between the chat window. If the chat window makes me twice as productive the command line tool probably makes me 5x as productive. At that kind of output on a developer salary the $70-200 a month is absolute peanuts compared to what you get in return 1: https://github.com/paul-gauthier/aider https://github.com/paul-gauthier/aider
- easygenes 3y agoFor what tasks do you consider 70B beyond GPT-3.5 performance? There are some I’m aware of, but they are very much the exception and not the rule, even with the best 70B fine-tunes currently available.
- syntaxing 3y agoI mainly use 70B for “text QA” on files I find sensitive like personal documents. The answers have been very close to what I get if I use GPT-3 (langchain makes it easy to switch). Do you use the quantized version? If so, try running the full one on a A100.
- ozr 3y agoI run 70B very cheaply using serverless GPUs. I've had the best experience with Runpod, but there are a few other options out there for it as well.