5 ms·
RAG feels hacky to me. We’re coming up with these pseudo-technical solutions to help but really they should be solved at the level of the model by researchers.
by Satam 2y ago
RAG feels hacky to me. We’re coming up with these pseudo-technical solutions to help but really they should be solved at the level of the model by researchers. Until this is solved natively, the attempts will be hacky duct-taped solutions.
- repeekad 2y agoWhat about fresh data like an extremely relevant news headline that was published 10 minutes ago? Private data that I don’t want stored offsite but am okay trusting an enterprise no log api? Providing realtime context to LLMs isn’t “hacky”, model intelligence and RAG can complement each other and make advancements in tandem
- jstummbillig 2y agoI don't think the parents idea was to bake all information into the model, just that current RAG feels cumbersome to use (but then again, so do most things AI right now) and information access should be intrinsic part of the model.
- deleted 2y ago[deleted]
- randomdata 2y agoIs there a specific shortcoming of the model that could be improved, or are we simply seeking better APIs?
- PaulHoule 2y agoOne of my favorite cases is sports chat. I'd expect ChatGPT to be able to talk about sports legends but not be able to talk about a game that happened last weekend. Copilot usually does a good job because it can look up the game on Bing and them summarize but the other day i asked it "What happened last week in the NFL" and it told me about a Buffalo Bills game from last year (did it know I was in the Bills geography?) Some kind of incremental fine tuning is probably necessary to keep a model like ChatGPT up to date but I can't picture it happening each time something happens in the news.
- ec109685 2y agoFor the current game, it seems solvable by providing it the Boxscore and the radio commentary as context, perhaps with some additional data derived from recent games and news. I think you’d get a close approximation of speaking with someone who was watching the game with you.
- viraptor 2y agoThat's so vague I can't tell what you're suggesting. What specifically do you think needs solving at the model level? What should work differently?
- Satam 2y agoThere’s probably lack of cpabalities on multiple fronts. RAG might have the right general idea but currently the retrieval seems to be too seperated from the model itself. I don’t know how our brains do it, but retrieval looks to be more integrated there. Models currently also have no way to update themselves with new info besides us putting data into their context window. They don’t learn after the initial training. It seems if they could, say, read documentation and internalize it, the need for RAG or even large context windows would decrease. Humans somehow are able to build understanding of extensive topics with what feels to be a much shorter context-window.
- simonw 2y agoDon't forget the importance of data privacy. Updating a model with fresh information makes that information available to ALL users of that model. This often isn't what you want - you can run RAG against a user's private email to answer just their queries, without making that email "baked in" to the model.
- viraptor 2y agoYou don't need to update the whole model for everyone. Fine tuning exists and is even available as a service in openai. The updates are only visible in the specific models you see.
- simonw 2y agoMaintaining a fine-tuned model for every one of your users - even with techniques like LoRA - sounds complicated and expensive to me!
- 2y ago
- ac1spkrbox 2y agoThe set of techniques for retrieval is immature, but it's important to note that just relying on model context or few-shot prompting has many drawbacks. Perhaps the most important is that retrieval as a task should not rely on generative outputs.
- danielbln 2y agoIt's also subject to significantly more hallucination when the knowledge is baked into the model, vs being injected into the context at runtime.
- l72 2y agoI've described it this way to my colleagues: RAG is a bit like having a pretty smart person take an open book test on a subject they are not an expert in. If your book has a good chapter layout and index, you probably do an ok job trying to find relevant information, quickly read it, and try to come up with an answer. But your not going to be able to test for a deep understanding of the material. This person is going to struggle if each chapter/concept builds on the previous concept, as you can't just look up something in Chapter 10 and be able to understand it without understanding Chapter 1-9. Fine-tuning is a bit more like having someone go off and do a phd and specialize in a specific area. They get a much deeper understanding for the problem space and can conceptualize at a different level.
- CGamesPlay 2y agoWhat you said about RAG makes sense, but my understanding is that fine-tuning is actually not very good at getting deeper understanding out of LLMs. It's more useful for teaching general instructions like output format rather than teaching deep concepts like a new domain of science.
- bashfulpup 2y agoThis is true if you don't know what you're doing, so it is good advice for the vast majority. Fine tuning is just training. You can completely change the model if you want make learn anything you want. But there are MANY challenges in doing so.
- CGamesPlay 2y agoThis isn't true either, because if you don't have access to the original data set, the model will overfit on your fine tuning data set and (in the extreme cases) lose its ability to even do basic reasoning.
- bashfulpup 2y agoAgain, that's why I said it is challenging. I regularly do fine tuning on a model with fine results and little damage to the base functionality. It is possible, but it's too complex for the majority of users. It requires a lot of work per dataset you want trained on.
- williamtrask 2y agoFwiw, I used to think this way too but LLMs are more RAG-like internally than we initially realised. Attention is all you need ~= RAG is a big attention mechanism. Models have reverse curse, memorisation issues etc. I personally think of LLMs as a kind of decomposed RAG. Check out DeepMind’s RETRO paper for an even closer integration.
- jejeyyy77 2y agoThe biggest problem with RAG is that the bottleneck for your product is now the RAG (i.e, results are only as good as what your vector store sends to the LLM). This is a step backwards. Source: built a few products using RAG+LLM products.
- zby 2y agoI guess you can imagine an LLM that contains all information there is - but it would have to be at least as big as all information there is or it would have to hallucinate. And also you Not to mention that it seems that you would also require it to learn everything immediately. I don't see any realistic way to reach that goal. To reach their potential LLMs need to know how to use external sources. Update: After some more thinking - if you required it to know information about itself - then this would lead to some paradox - I am sure.
- bashfulpup 2y agoA CL agent is next generation AI. When CL is properly implemented in an LLM agent format, most of these systems vanish.