8 ms·
The RAG Obituary: Killed by agents, buried by context windows
- deleted 1y ago[deleted]
- thenewwazoo 1y ago[flagged]
- sebmellen 1y agoIt truly is unfortunate. Thankfully most people seem to have an innate immune response to this kind of RLHF slop.
- Retr0id 1y agoUnfortunately this can't be true, otherwise it wouldn't be a product of RLHF.
- phainopepla2 1y agoCrowds can have terrible taste, even if they're made up of people with good (or at least middling) taste
- sebmellen 1y agoGo on an average college campus, and almost anyone can tell you when an essay was written with AI vs when it wasn't. Is this a skill issue? Are better prompters able to evade that innate immune response? Probably yes. But the revulsion is innate.
- deleted 1y ago[deleted]
- tptacek 1y agoThere are typos in it, too. I don't think this kind of style critique is really on topic for HN. Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- titanomachy 1y ago"This wasn't written by a person" isn't a tangential style critique.
- sebmellen 1y agoThose guidelines that you reference talk almost exclusively about annoyances on the webpage itself, not the content of the article. I think it's fair to point out that many articles today are essentially a little bit of a human wrapper around a core of ChatGPT content. Whether or not this was AI-generated, the tells of AI-written text are all throughout it. There are some people who have learned to write like the AI talks to them, which is really not much of an improvement over just using the AI as your word processor.
- bigwheels 1y agoDo you agree that bickering over AI-generated vs. not AI-generated makes for dull discussion? Sliding sewing needles deep into my fingernail bed sounds more appealing than nagging over such minutiae.
- xgulfie 1y ago[flagged]
- dymk 1y agoI don’t mind articles that have a hint of “an AI helped write this” as long as the content is actually informationally dense and well explained. But this article is an obvious ad, has almost no interesting information or summaries or insights, and has the… weirdly chipper? tone that AI loves to glaze readers with.
- tptacek 1y agoHow is this an ad? It's a couple thousand words about how they built something complicated that was then obsoleted.
- serf 1y agoin the same vein that a 'Behind The Scenes Look At The Making of Jurassic Park' is , in fact, an ad. having a company name pitched at you within the first two sentences is a pretty good give away.
- tptacek 1y ago3/4 of what hits the front page is an "ad" by that standard. I don't see how you can get less promotional than a long-form piece about why your tech is obsolete. Seems just mean-spirited.
- dymk 1y agoIt’s because the article’s main goal is to sell me the company’s product, not inform me about RAG. It’s a zero calorie article.
- nbstme 1y agohaha so true!
- SV_BubbleTime 1y ago> 3/4 of what hits the front page is an "ad" by that standard. Is anyone disagreeing with that?
- momojo 1y agoI'm guessing first draft was AI. I had to re-read that part a couple times because the flow was off. That second paragraph was completely unnecessary too since the previous paragraph already got the point across that "context window small in 2022". On the whole though, I still learned a lot.
- nbstme 1y agoThanks! Sorry if the flow was off
- tomhow 1y agoWe've been asking people not to comment like this on HN. We can never know exactly how much an individual's writing is LLM-generated, and the negative consequences of a false accusation outweigh the positive consequences of a valid one. We don't want LLM-generated content on HN, but we also don't want a substantial portion of any thread being devoted to meta-discussion about whether a post is LLM-generated, and the merits of discussing whether a post is LLM-generated, etc. This all belongs in the generic tangent category that we're explicitly trying to avoid here. If you suspect it, please use the the established approaches for reacting to inappropriate content: if it's bad content for HN, flag it; if it's a bad comment downvote it; and if there's evidence that it's LLM-generated, email us to point it out. We'll investigate it the same way we do when there are accusations of shilling etc, and we'll take the appropriate action. This way we can cut down on repetitive, generic tangents, and unfair accusations.
- themanmaran 1y agoI'm always amazed at claude codes ability to build context by just putting grep in a for loop. It's pretty much the same process I would use in an unfamiliar code base. Just ctrl+f the file system till I find the right starting point.
- eru 1y agoThat's what I used to use as a human, but then I finally overcame my laziness in setting up integration between my editor and compiler (and similar) and got 'jump to definition' working. (Well, I didn't overcome my laziness directly. I just switched from being lazy and not setting up vim and Emacs with the integrations, to trying out vscode where this was trivial or already built in.)
- lukaslalinsky 1y agoDo you trust 'jump to definition'. Obviously it depends on the language server, but it's best effort. I'm often frustrated when it doesn't work, because I broke the code in some way. Or it jumps to a specific definition, but there are multiple. If I was as quick at opening and reading files as claude code, I'd prefer grep with context around the searched term.
- eru 1y ago> Do you trust 'jump to definition'. It depends, for some languages 'jump to definition' tools ask the same compiler/interpreter that you use to build your code, so it's as accurate as it gets, and it's not 'best effort'. It also depends a bit on your project, some project are more prone to re-using names or symbols. > If I was as quick at opening and reading files as claude code, I'd prefer grep with context around the searched term. Well, Claude probably also doesn't want to have to 'learn' how to use all kinds of different tools for different languages and eco-systems.
- guipsp 1y agoIn java, for example, jump to definition is pretty flawless.
- kingjimmy 1y agoHas it not dawned on the author how ironic calling embeddings and retrieval pipelines "a nightmare of edge cases" when talking about LLM
- nbstme 1y agoHaha! LLMs themselves are pure edge cases because they are non-deterministic. But if you add a 7-step pipeline on top of that, it's edge cases on top of edge cases.
- djoldman 1y ago... for this specific use case (financial documents). These corpora have a high degree of semantic ambiguity among other tricky and difficult to alleviate issues. Other types of text are far more amenable to RAG and some are large enough that RAG will probably be the best approach for a good while. For example: maintenance manuals and regulation compendiums.
- nbstme 1y agoWhy? What if LLMs could parallelize much of their reading and then summarize the findings into a markdown file, eliminating the need for complicated search?
- redwood 1y agoWeird to see the use case referenced specifically code search when that's a very targeted one rather than what general purpose agents (or RAG) use cases might target.
- aussieguy1234 1y agogrep was invented at a time when computers had very small amounts of memory, so small that you might not even be able to load a full text file. So you had tools that would edit one line at a time, or search through a text file one line at a time. LLMs have a similar issue with their context windows. Go back to GPT-2 and you wouldn't have been able to load a text file into its memory. Slowly the memory is increasing, same as it did for the early computers.
- nbstme 1y agoAgree. It's a context/memory issue. Soon LLMs will have a 10M context window and they won't need to search. Most codebases are less than 10M tokens.
- pcthrowaway 1y agoWhen dependencies are factored in, I don't know if this is true.
- selcuka 1y agoI don't find this surprising. We are constantly finding workarounds for technical limitations, then ditch them when the limitation no longer exists. We will probably be saying the same thing for LLMs in a few years (when a new machine learning related TLA becomes the hype).
- nbstme 1y ago100%. The speed of change is wild. With each new model, we end up deleting thousands of lines of code (old scaffolding we built to patch the models’ failures.)
- sergiotapia 1y ago>The winners will not be the ones who maintain the biggest vector databases, but the ones who design the smartest agents to traverse abundant context and connect meaning across documents. So if one were building say a memory system for an AI chat bot, how would you save all the data related to a user? Mother's name, favorite meals, allergies? If not a Vector database like pinecone, then what? Just a big .txt file per user?
- nbstme 1y agoExactly. Just a markdown file per user. Anthropic recommends that.
- queenkjuul 1y agoAny kind of database is far too efficient for an LLM, just take all your markdown and turn it into less markdown.
- ako 1y agoThat is what Claude Sonnet 4.5 is doing: https://youtu.be/pidnIHdA1Y8?si=GqNEYBFyF-3Klh4- https://youtu.be/pidnIHdA1Y8?si=GqNEYBFyF-3Klh4-
- davidmckayv 1y agoThis glosses over a fundamental scaling problem that undermines the entire argument. The author's main example is Claude Code searching through local codebases with grep and ripgrep, then extrapolates this to claim RAG is dead for all document retrieval. That's a massive logical leap. Grep works great when you have thousands of files on a local filesystem that you can scan in milliseconds. But most enterprise RAG use cases involve millions of documents across distributed systems. Even with 2M token context windows, you can't fit an entire enterprise knowledge base into context. The author acknowledges this briefly ("might still use hybrid search") but then continues arguing RAG is obsolete. The bigger issue is semantic understanding. Grep does exact keyword matching. If a user searches for "revenue growth drivers" and the document discusses "factors contributing to increased sales," grep returns nothing. This is the vocabulary mismatch problem that embeddings actually solve. The author spent half the article complaining about RAG's limitations with this exact scenario (his $5.1B litigation example), then proposes grep as the solution, which would perform even worse. Also, the claim that "agentic search" replaces RAG is misleading. Recent research shows agentic RAG systems embed agents INTO the RAG pipeline to improve retrieval, they don't replace chunking and embeddings. LlamaIndex's "agentic retrieval" still uses vector databases and hybrid search, just with smarter routing. Context windows are impressive, but they're not magic. The article reads like someone who solved a specific problem (code search) and declared victory over a much broader domain.
- CuriouslyC 1y agoAgentic retrieval is really more a form of deep research (from a product standpoint there is very little difference). The key is that LLMs > rerankers, at least when you're not at webscale where the cost differential is prohibitive.
- nbstme 1y agoLLMs > rerankers. Yes! I don't like rerankers. They are slow, the context window is small (4096 tokens), it's expensive... It's better when the LLM reads the whole file versus some top_chunks.
- 1y ago
- CuriouslyC 1y agoRAG isn't dead, RAG is just fiddly, you need to tune retrieval to the task. Also, grep is a form of RAG, it just doesn't use embeddings.
- nbstme 1y agoYes my point is that the entire RAG pipeline like ingest, chunk, embed, search with Elastic, rerank is in decline. Grep is far simpler. It’s trivial.
- tw1984 1y agoNo, grep is not RAG. RAG is all about embeddings + vector search + LLM working under a fixed workflow. Saying grep is also RAG is like saying ext4 + grep is a database.
- rpcorb 1y agoYou don't decide what "RAG" means, a term that's been around for much, much less time than "database".
- CuriouslyC 1y agoSo you're saying grep isn't a form of information retrieval?
- tw1984 1y agoinformation retrieval is a much larger superset of RAG. grep + agentic LLM is not RAG.
- leobg 1y agoRetrieval Augmented Generation. How you do the retrieving is irrelevant. You can do it manually and it’ll still be RAG. Also, most RAG pipelines combine multiple approaches - BM25, embeddings, etc..
- deleted 1y ago[deleted]
- dkga 1y agoRAG is the new US dollar, now every year someone will predict its looming death…
- nbstme 1y agoHAHAHA. Ok let's call it "transformation." As i wrote "The next decade of AI search will belong to systems that read and reason end-to-end. Retrieval isn’t dead—it’s just been demoted."
- jgalt212 1y agoI'm not feeling it. Constantly pinging these yuge LLMs is not economic and not good for sensitive docs.
- nbstme 1y agoBut don’t you think LLM pricing is heading toward zero? It seems to halve every six months. And on privacy, you can hope model providers won’t train on your data, (but there’s no guarantee)
- queenkjuul 1y agoI don't see how it can trend to zero when none of the vendors are profitable. Uber and doordash et. al. increased in price over time. The era of "free" LLM usage can't be permanent
- dangoodmanUT 1y agoGoogle’s inference is profitable
- jgalt212 1y agoNot on the SERP page. The zero click Internet is bad for content producers and for those who sell ads (Google).
- imiric 1y agoOh, it's going to be "free" alright, in the same way that most web services are today. I.e., you will pay for it with your data and attention. The only difference is that the advertising will be much more insidious and manipulative, the data collection far easier since people are already willingly giving it up, and the business much more profitable. I can hardly wait.
- cmenge 1y agoWe're processing tenders for the construction industry - this comes with a 'free' bucket sort from the start, namely that people practically always operate only on a single tender. Still, that single tender can be on the order of a billion tokens. Even if the LLM supported that insane context window, it's roughly 4GB that need to be moved and with current LLM prices, inference would be thousands of dollars. I detailed this a bit more at https://www.tenderstrike.com/en/blog/billion-token-tender-rag/ https://www.tenderstrike.com/en/blog/billion-token-tender-ra... And that's just one (though granted, a very large) tender. For the corpus of a larger company, you'd probably be looking at trillions of tokens. While I agree that delivering tiny, chopped up parts of context to the LLM might not be a good strategy anymore, sending thousands of ultimately irrelevant pages isn't either, and embeddings definitely give you a much superior search experience compared to (only) classic BM25 text search.
- elliotto 1y agoI work at an AI startup, and we've explored a solution where we preprocess documents to make a short summary of each document, then provide these summaries with a tool call instruction to the bot so it can decide which document is relevant. This seems to scale to a few hundred documents of 100k-1m tokens, but then we run into issues with context window size and rot. I've thought about extending this as a tree based structure, kind of like an LLM file system, but have other priorities at the moment. Embeddings had some context size limitations in our case - we were looking at large technical manuals. Gemini was the first to have a 1m context window, but for some reason its embedding window is tiny. I suspect the embeddings might start to break down when there's too much information.
- codyb 1y agoFor anyone unfamiliar, construction tenders are part of the project bidding process and appear to be a structured and formal manner in which contractors submit bids for large projects.
- catlover76 1y agoThe Agents are just RAG
- intalentive 1y agoAgentic search with a handful of basic tools (drawn from BM25, semantic search, tags, SQL, knowledge graph, and a handful of custom retrieval functions) blows the lid off RAG in my experience. The downside is it takes longer. A single “investigation” can easily use 20-30 different function calls. RAG is like a static one-shot version of this and while the results are inferior the process is also a lot faster.
- nsomaru 1y agoHey, I’m interested in what you call “agentic search”. Did you roll your own or are you using a set of integrated tools? I’ve used LightRAG and looking to integrate it with OpenWebUI and possibly air weave which was a show HN earlier. My data is highly structured and has references between documents, so I wanted to leverage that structure for better retrieval and reasoning.
- intalentive 1y agoRolled my own in Python. For graph/tree document representations, it’s common in RAG to use summaries and aggregation. For example, the search yields a match on a chunk, but you want to include context from adjacent chunks — either laterally, in the same document section, or vertically, going up a level to include the title and summary of the parent node. How you integrate and aggregate the surrounding context is up to you. Different RAG systems handle it differently, each with its own trade offs. The point is that the system is static and hardcoded. The agentic approach is: instead of trying to synthesize and rank/re-rank your search results into a single deliverable, why not leave that to the LLM, which can dynamically traverse your data. For a document tree, I would try exposing the tree structure to the LLM. Return the result with pointers to relevant neighbor nodes, each with a short description. Then the LLM can decide, based on what it finds, to run a new search or explore local nodes.
- mscbuck 1y agoI've found his hybrid approach pretty good for the majority of use cases. BM25 (maybe Splade if you want a blend of BOW/Keyword), + Vectors + RRF + re-rank works pretty damn well. The trick that has elevated RAG, at least for my use cases, has been having different representations of your documents, as well as sending multiple permutations of the input query. Do as much as you can in the VectorDB for speed. I'll sometimes have 10-11 different "batched" calls to our vectorDB that are lightning quick. Then also being smart about what payloads I'm actually pulling so that if I do use the LLM to re-rank in the end, I'm not blowing up the context. TLDR: Yes, you actually do have to put in significant work to build an efficient RAG pipeline, but that's fine and probably should be expected. And I don't think we are in a world yet where we can just "assume" that large context windows will be viable for really precise work, or that costs will drop to 0 anytime soon for those context windows.
- kixiQu 1y agoThis is a great example of a piece with enough meaningful and useful content in it that it's very clear the author had something of value to deliver, and I'm grateful for that... but enough repetitive LLM-output that I'm very annoyed by the end. Actually, let me be specific: everything from "The Rise of Retrieval-Augmented Generation" up to "The Fundamental Limitations of RAG for Complex Documents" is good and fine as given, then from "The Emergence of Agentic Search - A New Paradigm" to "The Claude Code Insight: Why Context Changes Everything" (okay, so the tone of these generated headings is cringey but not entirely beyond the pale) is also workable. Everything else should have been cut. The last four paragraphs are embarrassing and I really want to caution non-native English speakers: you may not intuitively pick up on the associations that your reader has built with this loudly LLM prose style, but they're closer to quotidian versions of the [NYT] delusion reporting than you likely mean to associate with your ideas. [NYT]: https://www.nytimes.com/2025/08/08/technology/ai-chatbots-delusions-chatgpt.html https://www.nytimes.com/2025/08/08/technology/ai-chatbots-de...
- OutOfHere 1y agoThat's quite the over-generalization. RAG fundamentally is: topic -> search -> context -> output. Agents can enhance it by iterating in a loop, but what's inside the loop is not going away.
- cyberax 1y agoI wonder if something like LSP or IntelliJ's reverse index would work better for AI than RAG.
- devmor 1y agoThis reads like someone AI-generated prose to defend something they want to invest in and decry something it competes with. It does not come off as honest, written by a human, or useful to anyone outside of the specific, narrow contexts the "author" sees for the technologies mentioned. Frankly, reading through this at makes me feel as though I am a business analyst or engineering manager being presented with a project proposal from someone very worried that a competing proposal will take away their chance to shine. As it reaches the end, I feel like I'm reading the same thing, but presented to a Buzzfeed reader.
- cantor_S_drug 1y agoHow come this isn't the top comment? This post screams AI.
- jimbohn 1y agoFeels like saying Elasticsearch (and similar) tools are dead because we can just grep our way through things. I'd love to see more data on this.
- sublimefire 1y agoSaying that RAG alone is complex and should be superseded by agentic search is a bit weak. Agentic search makes more sense when your pipeline becomes more complicated: RAG+MCP+Client calls, it is then when you can see that LLM starts behaving erratically and cannot answer the question well. You then want better control over streams of content and intents which could be solved by smaller agents looping over the data.
- masterkram 1y agoAfter building a few RAG based apps I was curious to try the ClaudeCode based approach that is mentioned by the author. So I built a python service that exposes ripgrep to a rest api: https://github.com/masterkram/jaguar https://github.com/masterkram/jaguar This makes it possible to quickly deploy this on coolify and quickly build an agent that can use ripgrep on any of your uploaded files.
- zwaps 1y agoI am so tired of these undifferentiated takes. These types of articles regularly come from people who don't actually build SCALE systems with LLMs. Or, people who want to sell you on a new tech. And the frustrating thing is: They ain't even wrong. Top-K RAG via vector search is not a sufficient solution. It never really was for most interesting use-cases. Of course, take easiest and most structured - in a sense perfectly indexed - data (code repos) and claim that "RAG is dead". Again. Now try this with billions of unstructured tokens where the LLM really needs to do something with the entire context (like, confirm that something is NOT in the documents), where even the best LLM loses context coherence after like 64k tokens for complex tasks. Good luck! The truth is: Whether its Agentic RAG, Graph RAG, or a combination of these with ye olde top-k RAG - it's still RAG. You are going to Retrieve, and then you are going to use a system of LLM agents to generate stuff with it. You may now be able to do the first step smarter. It's still Rag tho. The latest Antrophic whoopsy showed that they also haven't solved the context rot issue. Yes you can get a 1M context scaled version of Claude, but then the small/detail scale performance is so garbage that misrouted customers loose their effin mind. "My LLM is just gonna ripgrep through millions of technical doc pdfs identified only via undecipherable number-based filenames and inconsistent folder structures" lol, and also, lmao
- gengstrand 1y agoI agree. Permit me to rephrase. From this learning adventure https://www.infoq.com/articles/architecting-rag-pipeline/ https://www.infoq.com/articles/architecting-rag-pipeline/ I came to understand what many now call context rot. If you want quality answers, you still need relevance reranking and filtering no matter how big your context window becomes. Whether that happens in a search that is upfront in a one shot prompt or iteratively in a long session through an agentic system is merely an implementation detail.
- sakoht 1y agoPeople say “agents not RAG”, but one framing is that this describes RAG where the database is a file system and bash is the query language (with other cli tools installed it can use, including curl, jq, grep). With writing its own notes on the filesystem structure and maintaining them as a way to “index the database” It is still using code to selectively grab the chunks of data it needs rather than than putting everything in context. It’s just better RAG?
- alastairr 1y agoIsn't 'agentic search' just another form of RAG? information still gets retrieved and added to the prompt, even if the 'prompt' is levels down in the product and not visible to the user.
- Imanari 1y agoI get the reasoning behind “letting the agent use grep in a loop“, after all, it is very similar to how humans would explore a document base with ctrl-f. But wouldn’t humans also use vector search all the time if it were as available as ctrl-f? So maybe not ditch vector search but provide it as a tool to the agent. Increased complexity aside, letting the agent explore a huge document base with “vector search in a loop“ should be more powerful that with grep in a loop. Overall I liked the article.
- regularfry 1y agoMy mental model (in the "all models are wrong, some are useful" sense) is that vector search is the thing that gives you the terms to grep for.
- te_chris 1y agoOr the thing that ranks the term based result. That’s the fun these days: it’s all whatever fits your problem.
- Imanari 1y agoI like it. With this approach it feels like you also don’t need to fiddle as much with the details of your vector search and DB as that portion just gets you going and the actual retrieval happens with grep in a loop.
- alansaber 1y agoConceptually yes but practically speaking vector search results are generally not sufficiently good
- stoneyhrm1 1y agoI'm free to be corrected because I'm no expert in the field but isn't RAG just enriching context, it doesn't have to be semantic search, it could be an API call or grabbing info from a database.
- msukhareva 1y agoRAG was always somewhat of a Frankenstein combining two things that should not be combined: information retrieval based on string matching enhanced with embeddings and LLM that needs not string matching but informative texts. If string matching is good but information is poor or provides wrong context, it would only enforce hallucinations. Search, tool calling and connections should be a part of the system and trained together with LLM
- anshumankmr 1y agoHonestly, I am not sure when RAG had its heyday to be deserving an obituary. I still think that it is in its early days.
- athrowaway3z 1y agoI had seen RAG mentioned a lot before I had gotten into LLM agents. I assumed it was tricky and required real deep model training knowledge. My first real AI use (beyond copy-paste ChatGPT) was Claude Code. I figured out in a few days to just write scripts and CLAUDE.md how to use them. For instance, one that prints comments and function names in a file is a few lines of python. MCP seemed like context bloat when a `tools/my-script -h` would put it in context on request. Eventually stumbled on some more RAG a few weeks later, so decided to read up on it and... what? That's it? A 'prelude function' to dump 'probably related' things into the context? It seems so obviously the wrong way to go from my perspective, so am I missing something here?
- beastman82 1y ago> No need for similarity when you can use exact matches This is a weakness, not a strength of agentic search
- malshe 1y ago> Table Integrity: Financial tables are never split—income statements, balance sheets, and cash flow statements remain atomic units with headers and data together In 10-k and 10-q often there are no table headers. This is particularly true for the consolidated notes to financial statements section. Standalone tables could be pretty much meaningless because you won't even know what they are reporting. For example, a table that simply mentions terms like beginning balance and ending balance can be reporting inventory, warranty, or short term debt. But the table does not mention these metrics at all and there are no headers. So I am curious to know how Fintool uses standalone tables. Do you retain text surrounding the tables in the same chunk as the table?
- deleted 1y ago[deleted]
- esafak 1y agoThis is just wrong. As many here have said, grep is RAG; just the most primitive kind. It means you miss typos, synonyms, semantic matches (e.g., "the payment service"), and AST matches. I have to deal with this when I use grep-based agents by handholding them and overpaying. grep is just something that enabled CLI-based tools to get to market faster. grep's dominance will fade as the landscape matures. The current pattern seems to be to outsource RAG to an MCP.
- leopoldj 1y agoThe author is conflating RAG with vector search. I think. One can use any and all available search mechanisms, SQL, graph db, regex, keyword and so on, for the retrieval part.
- rohansood15 1y agoI don't get why folks are so dismissive here. If you ever saw Claude Code/Codex use grep, you will find that it constructs complex queries that encompass a whole range of keywords which may not even be present in the original user query. So the 'semantic meaning' isn't actually lost. And nobody is putting an entire enterprise's knowledge base inside the context window. How many enterprise tasks are there that need referencing more that a dozen docs? And even those that do, can be broken down into sub-tasks of manageable size. Lastly, nobody here mentions how much of a pain it is to build, maintain and secure an enterprise vector database. People spend months cleaning the data, chunking and vectorizing it, only for newer versions of the same data making it redundant overnight. And good look recreating your entire permissioning and access control stack on top of the vector database you just created. The RAG obituary is a bit provocative, and maybe that's intentional. But it's surprising how negative/dismissive the reactions in this thread are.
- innagadadavida 1y agoThe article is not making a proper distinction of scale and is probably due to the small scale problem that they solved. What is small scale and <10K documents / files can be easily processed with grep, find etc. For something at larger scale >1M documents etc. you will need to use search engine technology. You can definitely do the same agent approach for the large scale problem - we essentially need search, look at the results and issue follow up queries to get documents of interest. All that said, for the types of problem the OP is solving, it might just be better to create a project in Claude/ChatGPT and throw in the files there and get done with it. That approach has been working for over 2 years now and is nothing new.
- maerch 1y ago> The agent follows references like a human analyst would. No chunks. No embeddings. No reranking. Just intelligent navigation. I think this sums it up well. Working with LLMs is already confusing and unpredictable. Adding a convoluted RAG pipeline (unless it is truly necessary because of context size limitations) only makes things worse compared to simply emulating what we would normally do.
- bze12 1y agoThis post was definitely written by an llm.
- findjashua 1y agoRAG != EBR
- diamondfist25 1y agowhats everyone's RAG pipeline? I was using qdrant, but im considering moving to OpenSearch since i want something more complete w/ a dashboard that i can muck around with
- kohlerm 1y agoCursor's search is still better(faster and cheaper) then Claude Code. I just did some tests. It looks like they do agentic searches with query rewriting.
- keeganpoppen 1y agoRAG isnt dead, it just isnt being used correctly. agents get the correct answer, yes, just not fast. RAG is a performance optimization (and cost), same as it always has been. as agents improve, the nature of what and how to RAG will change, but it has no reason to not exist, not the opposite.
- BrokenLButton 1y agoI worked at a startup that heavily relied RAG, and this article definitely articulates most of the same issues we ran into. I do think RAG still has its place but it is definitely becoming a harder to justify case.