3 ms·
RAG vs. Context-Window in GPT-4: accuracy, cost, & latency
- swiftlyTyped 3y agoTL;DR: (RAG + GPT-4) delivers superior performance, at 4% of the cost.
- parminder_88 3y agoGreat insights Atai
- theJoShPENNER 3y agoGreat analysis
- nathan_tarbert 3y agoThis is really interesting!
- igorkotua 3y agoGreat article, but what about open-source LLMs? Results will be the same?
- swiftlyTyped 3y agoIt'd be interesting to test and find out. As the article shows, there is some evidence that OpenAI may be using a new embeddings model under the hood of assistants retrieval. If they are, and if it's substantially better than the competition, then open-source RAG may lag for a while. -- But if they're just using ada v2 (or if the embeddings improvement is in cost, rather than performance), there should be tremendous potential for open-source models in this space. First of all, ada v2 is an aging model that has solid open-source competition. But more importantly, it seems the key is an LLM agent loop that can best make use of the RAG primitives. Intuitively, I'd expect open-source models to be smart enough for very good results in this domain.
- schitupolu 3y agoGreat Article !!
- calebrjohn 3y agoUsing the full context window each time is a great way to ensure your app is slow, expensive, and in-accurate. RAG is crucial, anyone thats built an AI app knows.