5 ms·
In terms of building something that's usable (considering cost, speed, scale, etc) if comparing an OpenAI API call to these, it's difficult for me to see a curr
by fareesh 3y ago
In terms of building something that's usable (considering cost, speed, scale, etc) if comparing an OpenAI API call to these, it's difficult for me to see a current path where these have any viable application outside some niche scenario.
From what I understand, even to run locally you/your team needs to be able to afford a machine with a 4090. These are super expensive in some countries.
I played around with the smaller Llama/Alpaca models and it wasn't really viable to build anything with.
Not really seeing a use-case for fine-tuning either compared to just few-shot prompting.
Can someone fill me in on what I'm missing? It feels like I'm out of the loop
- smoldesu 3y agoI'm running Vicuna on a free 4core Oracle VPS, and it's perfectly usable for a Discord bot. Responses rarely take more than 15 seconds with <256 max token limit, and the responses are much more entertaining than GPT 3.5. I'm not using the streaming API my server software[0] offers, but if I did it would probably load somewhere between the speeds of GPT-3.5 and GPT-4. It's more or less the same time a human would take to compose the same message. So... not exactly a serious use-case. But it's what I'm using, and now I'm saving 10s of dollars on inferencing costs per month! [0] https://github.com/go-skynet/LocalAI https://github.com/go-skynet/LocalAI I'm also using this to improve acceleration - https://cloudmarketplace.oracle.com/marketplace/en_US/adf.task-flow?tabName=O&adf.tfDoc=%2FWEB-INF%2Ftaskflow%2Fadhtf.xml&application_id=125935163&adf.tfId=adhtf https://cloudmarketplace.oracle.com/marketplace/en_US/adf.ta...
- paulgb 3y agoNow I’m curious what your bot does!
- smoldesu 3y agoIt's one of those "say a keyword with a question, get a response" type bots. I added in a few other "prompt sources" though, where it grabs the first part of an RSS entry or HN comment and tries to autocomplete the rest. Mostly just a boring testbed for me to play with models, for free, with friends.
- fareesh 3y agoTIL Oracle has VPS offerings with a free tier. Are they any good? Is the free-tier time limited? This use-case is alright for a toy I guess - which is the extent that I was originally expecting these things to be useful for.
- smoldesu 3y agoThey're okay. This isn't the place for a full review of their offerings (especially considering everyone's mixed feelings on Oracle), but I'm confident that it's better than most 1core/$5 deals you'll find elsewhere. > Are they any good? Yep, free tier allows you to spec up to 24gb of RAM without paying, which is cool. The bottleneck is really the disk speed, but that's not an issue with mmaped models. There's enough cached memory that it loads instantly, so it's good-ish for this use case. > Is the free-tier time limited? No, but there are a lot of strings attached: - The cores are vCPUs, not dedi (duh) - You can't create new instances when demand is high (unless you add a credit card) - Technically Oracle reserves the right to shut down the instance if demand gets really high (although I haven't heard any stories about this personally) Proceed with caution. It's still a great place to start before you shell out $1/hr for dedi GPU rackspace.
- tylergetsay 3y agoThey will shut down the VPS if there's no activity on it, not sure how they detect this though
- ukuina 3y agoCPU idle % over a rolling time window (say, 15 minutes).
- anon373839 3y agoFine-tuning is a much better proposition than you’re giving it credit for. Papers are coming out demonstrating that 7B parameter models can outperform GPT-4’s quality when trained on a limited set of tasks. Yet, a 7B model offers comparatively cheap and fast inference. Furthermore, for a lot of use cases, few-shot prompting is infeasible because you need to supply 2-3k tokens worth of few-shot examples with every prompt in order to fully specify the behavior you want. (As an example, think of long-form summarization where you want the summary to adhere to certain rules.)
- sgu999 3y agoWhat kind of tasks? Could you give some links to the papers you're referring to?
- anon373839 3y agoSure, here are two: 1. Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks https://huggingface.co/papers/2305.14201 https://huggingface.co/papers/2305.14201 2. Gorilla: Large Language Model Connected with Massive APIs https://arxiv.org/abs/2305.15334 https://arxiv.org/abs/2305.15334 Consider also these 2 papers supporting the feasibility of fine-tuning: 3. LIMA: Less Is More for Alignment [showing that a very small number of high quality examples is sufficient to align a base model] https://arxiv.org/abs/2305.11206 https://arxiv.org/abs/2305.11206 4. QLoRA: Efficient Finetuning of Quantized LLMs [showing that LLMs can now be fine-tuned quickly on consumer-grade GPUs] https://arxiv.org/abs/2305.14314 https://arxiv.org/abs/2305.14314 —- Adding up these developments (all of which occurred during the span of one week), I don’t see how huge, slow, general-purpose models maintain their relevance in the long term, when a lean, domain-focused model is right there within reach of every application developer.
- fareesh 3y agoSorry I should have been more specific - I was limiting my question to the bigger models. The smaller (~7B) models are feasible with these approaches.
- senttoschool 3y ago>From what I understand, even to run locally you/your team needs to be able to afford a machine with a 4090. These are super expensive in some countries. That's because we're only half a year into LLMs becoming mainstream. Give it 3-4 years. The advancements in bringing down model size, optimizations, and newer GPUs, SoCs from Nvidia, AMD, Apple, Intel, Qualcomm, etc will make it so that top LLMs will run on a highend laptop/desktop.
- Zemtomo 3y agoThis is bleeding edge stuff. All advances in this direction do indicate that it will be easier and easier for more people to do things with it. This doesn't need to work for everyone. A 4090 costs today 2k, the 3090 with also 24gb costs today 1k and costed 2k.
- caeril 3y agoAny particular reason not to run this model on a single Jetson AGX Orin 64GB? GP Core count is much lower than than the 4090 but it still does 275 int8 TOPS for only $2k