4 ms·
Pruning RAG context down to what the answer actually needs
- rooftopzen 3mo agoCliche topic - from a few years ago (the "RAG is dead" vs "All You Need Is Advanced RAG" BS - it came in waves and cycles, spread by bots on social media networks). "Pruning RAG Context" is trying to recycle the old stuff (again), presuming the reader is naive (implies kapa.ai is not going anywhere). The current cycles were "openclaw" (I think that died), now we are on "harnesses" - when that dies the paid social media bots will give you something else. Shell game. Just declare / define dictionary as a variable in your prompt to carry forward (when you decide to continue using LLMs for certain things). Also either summarize or truncate history. 3-4 year old concept. Not a big thing.
- timcobb 3mo agoShallow dismissal
- wolvoleo 3mo ago"RAG is dead" was declared because of ever expanding context windows. However, resources consumed expand so much with context that I think there is a huge practical barrier there. In the beginning AI services were basically free so nobody cared but this is rapidly shrinking. Personally I see more in a combination of RAG and live querying during the thinking process (e.g. by tools). Also I don't think dumping any context that might be relevant into the model really helps accuracy. In my experience models just get lost when they get an overload of irrelevant stuff in their context and start overlooking the relevant parts even if it does fit the window.
- agentdev001 3mo agoAm I wrong to be somewhat peeved by the use of "RAG" in these contexts? I always read things like this, and wonder if instead the author should be saying "Semantic Retrieval" or something something Vector, etc. Retrieval augmented generation captures tool-use, and; semantic search of course is really just a tool under the hood. To make an anology, in my mind, this is akin to saying "fuel air mixture system" when referring to direct fuel injection specifically, when of course, a carburetor also lives in that category.
- tingletech 3mo agoI read “RAG Context” as “the retrieved content injected into the context window”
- 0x696C6961 3mo agoSo when an agent does "cat file.txt" that's RAG to you?
- EagnaIonat 3mo agoThe answer would be yes. It's about using stored knowledge to increase the accuracy of the answer and evidence surfacing. It doesn't have to be a vector database. Kapa is one of the few companies doing RAG right.
- 0x696C6961 3mo agoYour definition is diluted to the point of being worthless.
- kristiandupont 3mo agoIt is to me. And I agree that the term is losing value because it's becoming ubiquitous but it's the differentiation from the first versions of ChatGPT etc., which were purely user input -> LLM -> output driven.
- tingletech 3mo agoIn the context of the title of the hn submission, I read “RAG Context” as “the retrieved content injected into the context window” .... "read" read /rε d/ not /riːd/ > So when an agent does "cat file.txt" that's RAG to you? No, that might be "RAG Context" though.
- red_hare 3mo agoI agree. We're seeing more variants of "RAG" that aren't semantic at all (e.g. coding agents or simple memory systems that feed summary indexes directly into context). I think, over time, it's going to become a SQL / NoSQL sort of divide. There will be the right kind of RAG for the job and lots of forcing the wrong kind because the developer doesn't understand the nuances.
- Technical_Plant 3mo ago[flagged]
- tangsoupgallery 3mo ago[flagged]
- esafak 3mo agotl,dr: They used a rubric to have the LLM grade the chunks on a Likert scale. I think this is a good way to coax numbers out of an LLM.
- siquick 3mo agoLikert scale just doesn't work in LLM evals. My idea of 3/5 is different from your 3 and definitely different from an non-deterministic system's 3.
- esafak 3mo agoThat's what the rubric is for; it reduces the problem to NLP, which is the LLM's forte. The more objective you can make your rubric, the better.
- petesergeant 3mo agoMy dissertation used a Likert scale with ChatGPT-3 era LLMs, and it was both internally consistent over 5-6 runs on a given statement, and consistent with human raters. I don't think you can bat it away as simply "doesn't work"
- StackOptimist 3mo ago[flagged]
- lbaltensp 3mo ago"relevance-to-the-query is exactly the signal that misses the buried caveat." You are exactly hitting the nail on the head. The underlaying problem of a pruner is to determine document relevance to a user query. And document relevance comes in different flavours: 1. Direct relevance: a document can be relevant to the user query because it directly answers it. 2. Subquery-Conditional Relevance: a document can be relevant because it answers a decomposed subquery of the user query. 3. Document-Interdependent Relevance: a document is relevant only because another document provides bridge context, domain knowledge, disambiguation, a definition or a constraint. This means the right question is therefore not "is this document relevant to the original query in isolation?", but: Given the original user query, the decomposed sub-queries, and all other retrieved documents, does this document contribute to form a complete and sufficient set of documents for answering the original query? This is exactly why we tuned our pruner at kapa based on recall against this Document Relevance Model.
- deleted 3mo ago[deleted]
- wolvoleo 3mo agoI think this is really where energy saving comes into play. Context is so incredibly energy and processing time sensitive.
- chimpanzee2 3mo agoi wouldn't trust the small llm there. it will be an intellectual bottleneck when it comes to processing the very information your just arduously fished out of the ocean!!
- kordlessagain 3mo agoI've had good success using a local model for preprocessing (tag extraction) and tuning the search parameters (hybrid search).
- synapsehire 3mo ago[flagged]
- grewil2 3mo ago"Three knobs matter:" I can't help reading articles with the radar on for signs of AI-generation nowadays. I have noticed that Claude sometimes uses the word knob for parameter, so here I get suspicious.
- petesergeant 3mo agoThe contribution here is the technique (which looks like a nice alternative to reranking) rather than the article or it being some kind of think piece, so does it really matter? I was able to skim it and go from "lol they reinvented re-ranking" to "oh, that's something more interesting, I should try this one day", and the writing didn't get in the way.
- daveKoala 3mo agoI am being constantly mocked for using there term 'knob fiddling', but I am a British child of the 70's :-)
- fuck_google 3mo ago[dead]
- Avery29 3mo agoBad or only loosely related context can make the final answer worse than having less context.
- tauridev 3mo ago[flagged]
- alansaber 3mo agoWould be cool to see a retrieval comparison to IE a claude code agent trace for the same query, even a cherry picked one.
- Imanari 3mo agoDid you experiment with the Pruner completely replacing the Reranker?
- timedude 3mo agoYep seems like the pruner and reranker could be combined. That is what i built.
- lbaltensp 3mo agoYes, the pruner operates on the top 15 chunks after reranking. reply
- deleted 3mo ago[deleted]
- lbaltensp 3mo agoReplacing the reranker with the pruner completely was not feasable since it would have to run on the 150-200 chunks we retrieve before reranking.
- Alan_JoshyMJ 3mo ago[dead]