8 ms·
It's been enough time since this leaked, so my question is why aren't there blog posts already of people blowing their $300 of starter credit with ${cloud_provi
by linearalgebra45 4y ago
It's been enough time since this leaked, so my question is why aren't there blog posts already of people blowing their $300 of starter credit with ${cloud_provider} on a few hours' experimentation running inference on this 65B model?
Edit: I read the linked README.
> I was impatient and curious to try to run 65B on an 8xA100 cluster
Well?
- v64 4y agoThe compute necessary to run 65B naively was only available on AWS (and perhaps Azure, I don't work with them) and the required instance types have been unavailable to the public recently (it seems everyone had the same idea to hop on this and try to run it). In my other post here [1], the memory requirements have been lowered through other work, and it should now be possible to run the 65B on a provider like CoreWeave. [1] https://news.ycombinator.com/item?id=35028738 https://news.ycombinator.com/item?id=35028738
- linearalgebra45 4y agoAre you sure about that? I can't remember where I saw the table of memory requirements, but I'm sure some of the larger instances here [1] will surely be able to cope (assuming they're available!) Oracle gives you a $300 free trial, which equates to running BM.GPU4.8 for over 10 hours - enough for a focused day of prompting [1] https://www.oracle.com/cloud/compute/gpu/ https://www.oracle.com/cloud/compute/gpu/
- v64 4y ago> Are you sure about that? I'm not. The only way to know it is to try :) thank you for the link!
- linearalgebra45 4y agoYou only get a single month-long window to spend the credit! And I'm sure not going to spend any of my own money on prompting experiments. I might be suffering from FOMO to some degree, I've just got to tell myself that this won't have been the only time model weights get leaked!
- mynameisvlad 4y ago> And I'm sure not going to spend any of my own money on prompting experiments. This certainly sounds a lot like whining that others aren’t doing the work you yourself don’t want to do.
- linearalgebra45 4y ago"prompting experiments" is just my use-case. According to v64 a lot of people have had the same idea of spinning up a trial instance to run inference, which is unsurprising. I'm not in a position to put in any meaningful work towards optimising this model for lower-end hardware, or working on the tooling/documentation/user experience.
- smoldesu 4y agoThanks for sharing it! I'm using their "Always Free" tier to host an Ampere-accelerated GPT-J chatbot right now. Works like a charm, and best of all, it's free!
- jocaal 4y agoI don't understand, the Ampere they refer to in their free tier are cpu's not gpu's. How did you manage to do that
- smoldesu 4y agoCustom PyTorch with on-chip acceleration: https://cloudmarketplace.oracle.com/marketplace/en_US/listing/125935163 https://cloudmarketplace.oracle.com/marketplace/en_US/listin... Not as fast as a GPU, but less than 5 seconds for a 250 token response is good enough for a Discord bot.
- nl 4y agoThis is the most interesting thing I've read in this thread. How have I never heard of this accelerator?!
- damascus 4y agoDo you have any code from your discord bot you're willing to share? I'd be happy to share back any updates I made to it. I've been wanting to play with this idea for a bit.
- deleted 4y ago[deleted]
- fswd 4y agoIf you actually try and do this, the sales people will stop you due to some internal rule. No GPUs on free credit. Unless the situation has changed of course..
- MacsHeadroom 4y agoI'm running LLaMA-65B on a single A100 80GB with 8bit quantization. $1.5/hr on vast.ai
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- linearalgebra45 4y agoWhat instance are you using?
- sillysaurusx 4y agoCareful though — we need to evaluate llama on its own merits. It’s easy to mess up the quantization in subtle ways, then conclude that the outputs aren’t great. So if you’re seeing poor results vs gpt-3, hold off judgement till people have had time to really make sure the quantized models are >97% the effectiveness of the original weights. That said, this is awesome — please share some outputs! What’s it like?
- MacsHeadroom 4y agoThe output is at least as good as davinci. I think some early results are using bad repetition penalty and/or temperature settings. I had to set both fairly high to get the best results. (Some people are also incorrectly comparing it to chatGPT/ChatGPT API which is not a good comparison. But that's a different problem.) I've had it translate, write poems, tell jokes, banter, write executable code. It does it all-- and all on a single card.
- data_maan 4y agoIs it just the RLHF training for the prompting that makes a difference, or are there also other, more tangible differences?
- ulnarkressty 4y agohttps://medium.com/@enryu9000/mini-post-first-look-at-llama-4403517d41a1 https://medium.com/@enryu9000/mini-post-first-look-at-llama-... *later edit - not the 65G model, but the smaller ones. Performance seems mixed at first glance, not really competitive with ChatGPT fwiw.
- linearalgebra45 4y ago> not the 65G model, but the smaller ones Haha, that's right! I saw that one too
- minxomat 4y ago> not really competitive with ChatGPT That's impossible to judge. LLama is a foundational model. It has received neither instructional fine tuning (davinci-3) nor RLHF (ChatGPT). It cannot be compared to these finetuned models without, well, finetuning.