5 ms·
I lead AI teams at my company. I've advised leadership against any kind of training / fine-tuning anything. We're not in the business of training models. We wi
by LASR 3y ago
I lead AI teams at my company. I've advised leadership against any kind of training / fine-tuning anything.
We're not in the business of training models. We will never be as good as OpenAI / Anthropic etc.
Where the real value in applications is smarter prompting techniques and RAG. There is a lot of room at the bottom in doing "dumb" things and simply feeding models with the right context to deliver customer value.
- elforce002 3y agoThis. I work in a startup and told upper management we need to keep focusing on ML models that bring tangible benefits to our customers and then, try to integrate LLMs into their current flow instead of pivoting completely to LLMs. It seems they valued the input and now we're going for a hybrid approach.
- jnwatson 3y agoIt is trivial to fine tune these days. RAG is already irrelevant with large context windows.
- Xenoamorphous 3y ago> RAG is already irrelevant with large context windows Just last Friday I took the contents of the 2024 folder of one of the teams at the company I work for, for which we use RAG at the moment. I dumped the text index, concatenated it and used Google’s API to return the token count, to see if it would fit in Gemini’s 1M context window; turned out it was 5.7M tokens. And that’s less than 3 months worth of documents for that team. So yeah RAG is not dead yet, although I do question its usefulness, but that’s a separate topic.
- greenavocado 3y agoDid I read this correctly? You uploaded millions of words of your company's internal communications to Google?
- Xenoamorphous 3y agoI did. But this is under an enterprise deal with them that warrants privacy, not the generally available stuff. OpenAI has similar arrangements (Enterprise ChatGPT) and MS Azure before them.
- simonw 3y agoCitation needed on "trivial to fine tune".
- jnwatson 3y agoCheck out Google's AI Studio. It makes it easy to fine tune. Disclaimer: I work for Google.
- danielmarkbruce 3y agoThere is no citation needed. It is indeed trivial to fine-tune. Doing a good job is another matter, but the claim is correct. Google around and find a blog post showing how. The claim that RAG is dead is obviously wrong.
- simonw 3y agoFor "citation needed", read "please link me to a blog post showing how, don't just tell me to Google for one". The internet is full of blog posts about this. That doesn't mean they're actually good - I'd love to be pointed at one that has proven itself useful for someone (and definitely isn't just LLM blog-spam). I don't care if it's trivial to fine-tune and get crap results - I care about fine-tuning where the result was worth the effort. For the record, my favourite guide to fine-tuning is the section of this Jeremy Howard video that shows how to train a text-to-SQL model: https://www.youtube.com/watch?v=jkrNMKz9pWU&t=4850s https://www.youtube.com/watch?v=jkrNMKz9pWU&t=4850s
- danielmarkbruce 3y agoIt's an internet forum, not an academic journal. Water tight arguments are not needed. If one wants to call bs, they can just do it, no need to dance around the topic by asking for a citation.
- simonw 3y agoOK, I call BS. Fine-tuning an LLM is not "trivial" - especially if you want to get useful results, as opposed to just being able to say "look, I fine-tuned an LLM".
- SgtBastard 3y agoA remarkable comment in that it is clear, confident and wrong. Fine-tunes lead to catastrophic forgetting. RAG is only irrelevant if you’re completely disinterested in cost and latency. We also don’t have enough data to gauge performance of models >200k context window size when reasoning over inputs of that size, much of which will be irrelevant to any particular user. Multiple random needles in haystack tests work flawlessly, but rarely applies to real world activity.
- seydor 3y agoYour advise is based on what?
- throwaway74432 3y agoHear hear. I know a 3 person startup that has a "lead AI researcher" who is trying to train and fine-tune models. That's not their startup's purpose though... they have an actual product. So wtf are they doing? The lead AI guy thinks he's going to compete with these big companies and it's total fantasy. LLMs are a commodity
- qeternity 3y agoThat does indeed sound crazy. But finetuning is also a commodity these days. You can train a good Mistral LoRA in under 24 hours on a single consumer GPU. We’re talking about $10 of compute. You can run a dozen of these LoRAs atop the same base model on the same infrastructure for a dozen specific use cases. The inference quality, performance and cost can all be substantially better than GPT4 with prompting.
- deleted 3y ago[deleted]
- VirusNewbie 3y agoDoesn't it entirely depend on how specialized the training data for a given fine tuned model might be?
- redox99 3y agoThat's a pretty odd stance. I've finetuned llama/mistral models that greatly outperform GPT4 with just a prompt. You have to know when to RAG, finetune, or RAG+finetune.
- kossTKR 3y agoHow narrow is the dataset to be outperforming greatly? Just curious about what the usecase is for a 7b model in a business context - ie. what does it do?
- redox99 3y agoCode assistant for a niche programming language that GPT4 knows very little about and barely gets a hello world right.
- singularity2001 3y agogreatly outperform GPT4 *for* just a prompt your overfitting to training data convinces no-one that you created a "better GPT4"
- redox99 3y agoDo you always assume other people are incompetent? That's not very nice of you. I mostly work on AI, so I know if I'm overfitting or not. It performs provably better in it's domain (a niche programming language). GPT4 can barely write a hello world for it. I'm not creating a "better GPT4" general chatbot. I'm finetuning for a specific task.
- lostmsu 3y agoYou are making an extraordinary claim, and they require extraordinary evidence. Unless presented it is a good idea to assume they are bogus.
- simonw 3y ago"I've finetuned llama/mistral models that greatly outperform GPT4 with just a prompt" If you write about your experiments with that in detail I guarantee you'll get a lot of interest. The community is crying out for good, well documented, replicable examples of this kind of thing.