8 ms·
RAG Is Simpler Than You Think
- 7734128 1mo agoThere have been many blogs like this over the last years. Yes, embeddings are computationally heavy, but they are not at all complicated and they provide a lot of benefit. 90% of "document" based RAG projects should view semantic search with embeddings as their primary method. It's very powerful and so easy to implement that you could try it out and discover whether performance would be an issue rather than trying to anticipate it.
- petesergeant 1mo agoEmbeddings are reasonably simple, but it’s a journey to get there, and I am very proud of the dog-heavy explainer I wrote on them: https://sgnt.ai/p/embeddings-explainer/ https://sgnt.ai/p/embeddings-explainer/
- dizhn 1mo agoThis is very good. Thanks.
- dotancohen 1mo agoThis is terrific, thank you! There's a typo in the following sentence: > we don’t especially want to say that books on forestry and similar to books on puppies ^and^are
- rglover 1mo agoStarted reading and will have to finish later but thank you for sharing. Very helpful post.
- pantsforbirds 1mo agoI think it's VERY project specific. If you are looking for anything technical at all, then keyword search almost always does better (in my experience). I'd actually recommend starting with keyword search, and then expanding with embeddings after you have a better idea of what your users are trying to determine.
- aitchnyu 1mo agoHow is pgvector with Sentence Transformers, a CPU-only embedding model, compared to a model hosted by OpenAI?
- cloudoora 1mo ago[dead]
- nilirl 1mo agoMaybe I'm old but where exactly are the "dragons"? How is RAG any different from the search systems we've been building before LLMs? Is it the sudden need for everyone to design a search API and engine that's driven this trend? If so, I'd like to see more design patterns around existing search problems: - Correcting or backtracking based on feedback. - Measuring relevance. - Comparison with task-based pre-written queries. Does every LLM task need a full blown search engine? Why not a tightly scoped domain API for data retrieval?
- TudorAndrei 1mo agoIt's just information retrieval packaged as something new.
- kachnuv_ocasek 1mo agoAnd you can't fundraise on some old "information retrieval".
- mdp2021 1mo agoIt's just information retrieval through a new NN based technology that allows to map concepts and ideas as the compression of long text into points in a multidimensional space that manages to compress even more dimensions than the given ones, through non-transparent engines that give different mappings and results, and still (the information retrieval) requires many more clever tricks than the simple idea of vector distance ordering because things do not quite work as they should. Let's say it's just "computation packaged as something new". "Trivial things".
- brabel 1mo agoThe whole embedding thing which converts “tokens” to vectors, which you then store in a vector database so that you can later query by vector distance, seems to be LLM specific technology, no? As far as I know the vectors look a lot like the weights in a LLM itself which is why the vector search also works with some level of intelligence.
- 1mo ago
- Angostura 1mo agoI have a particular antipathy for articles too lazy to spell out acronyms on first use. So: https://en.wikipedia.org/wiki/Retrieval-augmented_generation https://en.wikipedia.org/wiki/Retrieval-augmented_generation
- _joel 1mo agoFor those times you need to Red Amber Green your BM25
- dotancohen 1mo agoThe audience for this piece is already very familiar with RAG. I don't want articles discussing e.g. OLED screens telling me what the acronym is - that would be a sign that the article is far below the level that I need.
- vaylian 1mo agoA hyperlink to Wikipedia would have solved that issue.
- redsocksfan45 1mo ago[dead]
- Lorean1 1mo agoMaybe if a person can't even google RAG they are not the intended audience of that article.
- Zambyte 1mo agoEh, a healthy web is a web. I enjoy my preferred search engine, but surfing the web is becoming a lost medium.
- tux3 1mo agoHypermedia? In my hypertext markup language? That is so not Web 5.0. Best I can offer is a support widget that pops up and keeps trying to talk to you until you interract with it.
- apavlinovic 1mo agoThe article sounds like AI slop with some predictable tells like short punctual sentences, bizarre jargon, and titles like "Recipe 4: On-The-Fly Embedding (The Fresh Data Play)" Can we not reward junk like this? Most of the sentences are incomprehensible and provide zero actual argumentation, it's just a list of "whats" with no "whys"
- khalic 1mo ago> Why this is more flexible than embeddings Oh boy...
- refactor_master 1mo agoHere’s an even simpler take: just embed everything the first time, then track what was changed. Use a cheap model to summarize and clean up the documents/chats with summary and keywords. Unless you have entire libraries of books to embed it’s going to be a few hundred dollars of API calls. Then, throw it all in BigQuery. Handles all the vector stuff natively. Sprinkle an agentic bot UI thing on top to make it appear all-knowing and magical. I assume other vendors than Google have a similar batteries-included approach you can just plug in.
- cpursley 1mo agoYep, lock into some vendor from day 1. Great idea!
- usernametaken29 1mo ago> embed everything the first time This assumes your text is small. Try embedding pdf reports - though luck. It surely won’t fit into most embeddings. I can think of many more examples: books, news articles, medical reports, insurance claims etc. they’re all too big to “index it all at once”
- robrorcroptrer 1mo agoWhat about splitting bigger content into chunks before embedding?
- bob1029 1mo agoAgentic query rewrite on top of good old fashioned Lucene is the end game. This is effectively providing a lot of the same magic you get with the semantic approach. Allowing the agent to query the document store iteratively is where the capabilities become unbounded. Embeddings and semantic search add non determinism on top of non determinism. This seems fundamentally cursed. Lexical is much easier to control, iterate and debug. The tools are incredibly mature. Your users will probably prefer it as well.
- jrochkind1 1mo ago[flagged]
- pixelbro 1mo agoI've not seen such a clipped cadence out of an LLM. I would not automatically suspect the GP. Maybe there's better ways to spend your time?
- jrochkind1 1mo agoMaybe people are just learning to write in that style LLMs learned to write from statistical people? "is where the capabilities become unbounded" is a weird thing to say and not really true. "is the end game", "add non determinism on top of non determinism", there are a lot of AI-isms in this short comment. But it's possible people are just learning to write this way now, I am curious if that's so too! As far as uses of time, you are engaging in this dialog too, if you find it not a good way to spend time I recommend ceasing!
- bob1029 1mo agoIt would seem AI psychosis flows both ways.
- jankovicsandras 1mo agoIf someone has a Postgres db and want very simple RAG: https://github.com/jankovicsandras/plpgsql_bm25 https://github.com/jankovicsandras/plpgsql_bm25 BM25 search implemented in PL/pgSQL ( Unlicense / Public domain ) The repo includes also plpgsql_bm25rrf.sql : PL/pgSQL function for hybrid search ( plpgsql_bm25 + pgvector ) with Reciprocal Rank Fusion; and Jupyter notebook examples.
- simianwords 1mo agoOT but its interesting that none of the harnesses today use embeddings but just simple grep. I would not have predicted this
- imtringued 1mo agoOk? I'm not seeing how that is interesting, you're exclusively focusing on coding which requires precise substring locations. Google is basically almost entirely driven by embedding models now.
- simianwords 1mo agoAnd why do you think coding didn’t benefit from embeddings? It was attempted many times and the industry gave up. I find this interesting because practically no one is doing RAG on thier personal data which is something I wouldn’t have expected.
- owen-hill 1mo ago[flagged]
- marginalia_nu 1mo agoA lot of this is due the size of the corpus. Grep falls apart for severely underspecified queries, which is the difficult part of web search. For any given query in web search there can be several millions of candidate results. You can get good results with FTS as well, but just finding phrase matches is inadequate, you need more ranking signals to find relevant results. When Claude is looking for a function in your code base, it needs to sift through dozens of matches. This is not hard, and anything beyond grep is likely not worth the effort.
- anthonypasq 1mo agocursor still uses embeddings and theyve found it works better than just grep https://cursor.com/blog/semsearch https://cursor.com/blog/semsearch
- 1mo ago
- usernametaken29 1mo agoI worked on large scale RAG systems before and can say people vastly underestimate full text search and vastly overestimate embeddings. FTS is really easy, portable and scalable and gets you very far, the 80/20 rule applies. Embeddings appear to be nice and magic but when you really get into them you notice: semantic similarity isn’t as good as you think and certainly it won’t make everyone happy. You will inevitably end up having to re-embed more or different chunks of your text to accommodate more and more precise embedding search - at which point you’ll go the last mile and do reranking etc etc all the while having to support the operational burden of vector search. Then you turn around and build a search query with 500 keywords and sure it’s painful but it just works, accommodates all use cases, scales and is overall less annoying to maintain.
- clevergadget 1mo agoI don't know what level of quality is required for this site but RAG is trash its just trash. its magic beans.
- lacedeconstruct 1mo agoI thought text search was always the first thing you try, then fuzzy search, then you go for RAG
- ozim 1mo agoI think Bitwarden implemented some vector search in their password search feature ... totally annoying it gives me back all kinds of stuff that I don't care. I want fuzzy search like 95% of time and then I might consider having additional list of things that can be suggested by vector search.
- a1o 1mo agoA good UI could do these and also exact match, give some point system to the results, then order them and perhaps use a bold highlight to reflect what parts of the input query reflected in each result.
- 1mo ago
- jrochkind1 1mo agoMore LLM-generated text about LLMs. Is anyone else actually finding it harder and harder to read LLM generated text? I find it quite tiring, my brain just does not want to get through it.
- polynomial 1mo agoThe enshittification of the web, now powered by AI.
- allexander 1mo agoIn the same boat here.
- EGreg 1mo agoIt’s largely because LLMs are reaching for many different types of adjectives or verbs in the same sentence, in a jarring way. While embedding it in a confidently declarative sentence. Everything sounds like some profound insight, dialed to an 11, but written as poetry. Especially those headings. With the short sentences.
- allexander 1mo agoI have to agree with you. Yet it is tiring, people don't even try anymore.
- Planktonne 1mo agoYour brain is incredibly adept at pattern recognition; it doesn't focus on LLM-generated text for the same reason it doesn't stare at wallpaper. We've all learnt that it's not really communication, and so can be dispensed with.
- inigyou 1mo agoI'm Becoming AI-Blind: https://news.ycombinator.com/item?id=49386699 https://news.ycombinator.com/item?id=49386699
- jrochkind1 1mo ago
- jmutex 1mo agoChunk size matters way more than the retrieval model in my experience. Get that wrong and nothing else helps.
- esafak 1mo agoDon't leave us hanging! How do you set it?
- KaseyKim 1mo agoi want to ask that, if a user want to search sth, but he doesnt know the exact name(keywords), just some description. at this moment, whether the text serach fail?
- timedude 1mo agoText search is not ideal for that. I such cases embedding works better
- gabosarmiento 1mo agoI would like to see how each recipe performs against its corresponding evals. Some sort of ranking would be useful. Everyone keeps posting articles about how to implement RAG, but I also wonder why there isn’t some sort of skill to help people create a simple retrieval plan, starting with the retrieval methods and connecting them with evals. This could show whether they actually improve the result and make retrieval simpler for any agent, instead of making people start from zero.
- autogn0me 1mo agoIt seems not many RAG compare themselves across the same benchmarks. https://ggozad.github.io/haiku.rag/ https://ggozad.github.io/haiku.rag/ Does an ok job. The part I don’t see being discuss is the whole RL agents writing code to perform RAG queries. It’s one thing haiku-rag does that’s interesting and would like to know what other RAG have that agentic querying with benchmarks
- j0selit0 1mo agoauthor here - that's an amazing idea. would be an insanely large article though - maybe will write up a series
- tobin1994 1mo ago[dead]
- manganate06 1mo ago[flagged]
- luciana1u 1mo ago[flagged]
- deleted 1mo ago[deleted]
- jillesvangurp 1mo agoRAG is basically good old information retrieval with LLMs doing the querying. This can include vector search but it works without that as well. Treating vector search as magic pixie dust that makes search great without effort is not necessarily going to work that well. Also, it can add a lot of cost and complexity to the equation. And if not tuned properly, you don't necessarily get good results. The key thing with RAG is to get the right information in the context with as few queries as possible. That requires good recall (ensuring that if it is there it can be found with a reasonable query) and precision (ensuring the best stuff is on top and minimizing false positives). With search, and by extension RAG, the principle of shit in, shit out applies. Most of what search teams did before AI and RAG is still the best way to optimize the experience with RAG. And if you mess that up, search is not going to be working that well and no amount of AI can compensate for that or only at great cost in tokens and time. So, having an ETL pipeline to pre-process what you index, testing & benchmarking search quality, etc. are all helpful. The good news is that you don't need that much skills with agentic coding to build something half decent for this. This code almost writes itself. And even a little bit of effort on extracting structure before indexing can make a big difference.
- dmix 1mo ago> With search, and by extension RAG, the principle of shit in, shit out applies Similar to SEO on marketing pages, we started rewriting product docs around the idea that it will be consumed by a RAG. Mostly by putting a lot of focus on well structured headlines, thinking more carefully about technical terminology vs common human-language questions, occasionally using variations of keywords in the text, etc. This applies to pure LLM consumption too, not just hybrid search. Once you start tracking what users are asking you learn to adapt the documentation around it. And LLMs can also suggest improvements by comparing questions vs search results vs LLM responses.
- jillesvangurp 1mo agoIt's a start. Where it gets tricky is companies with years/decades of highly unstructured data, duplicated documents, obsolete or draft versions of those documents, etc. And where it gets more tricky if the data is spread all over the place in weird tools, databases, spreadsheets, etc. that has some structure but is maybe a bit inconsistent, incomplete, or not that well documented. If you flatten all that into plain text and then create embeddings, you are effectively throwing out the baby with the bathwater. But on the other hand if you put some effort into normalizing and extracting some structured meta data, you gain a flexibility to do more sophisticated querying that get you more precise results. You can of course try to fix things at the source, which is a valid thing but usually not that practical when you have a lot of data to worry about.
- sangwook 1mo agoIm curious whether the $10,000 figure includes unstated migration costs, since the raw embedding API cost under the earlier assumptions comes to $10.
- hizyyo 1mo ago[flagged]
- unicorn_platfor 1mo ago[dead]
- Otterly99 1mo agoAlthought I agree with the first point of the author that FTS is underrated in this new RAG-first framework, the whole article really hides all the problems with RAG-pipeline and kind of hand wave everything. If you are building a RAG pipeline for your company and are struggling like me, I would recommend this author that has whole series on entreprise documents (start with the one from May 22nd): https://towardsdatascience.com/author/angela.shi/page/4/ https://towardsdatascience.com/author/angela.shi/page/4/ Note: I am not the author, just got her article in my newsletter and found it useful.
- ufocia 1mo agoWow! Terrible layout. Shouldn't fully justify on a small screen.
- pioneerjeff 1mo agoWhat RAG means for AI is what a library means for human beings. It's necessary and would be good for you if you want to learn something systematically. But for most of the normal issues, we can not rely a lot on it.
- bewareofscams 1mo agoRAG is so 2024.
- trivet 1mo agoStart with BM25 and only add embeddings when keyword search actually fails you. Saves a lot of pain.
- ivansavz 1mo agoDoes anyone have experience using SMLs for RAG (either as query rewriter or as generator for the final answer)? I'd like to work with a corpus offline (internal university research data) and I'm hoping I can get everything done without the data leaving the premises. I guess the biggest bottleneck is going to be for the context window size which won't be able to fit too many result "hits." Any info or advice would be appreciated.
- respectattentio 1mo agoI believe embedding-based RAG, everybody is using, will end. As chips advance, you would use a big llm instead of word embedding for retrieval. It's much more accurate and extensive covering every topic. Still need ~2 years to be replaced.
- inigyou 1mo agoHow would you use a big LLM for retrieval?
- respectattentio 1mo agoAs simple as a prompting it with structural output or restrictions for your criteria. With agents, the prompting could be dynamic for maximum accuracy for every retrieval. This absolutely would beat the best of the best embedding-based RAG models. Nobody uses this now mainly due to speed. An llm retrieval would be 10x or more slower than embedding. You can try that now Take some failing cases or bad retrieval from your current system Prompt an llm wisely like a perfect prompt to get what you want and provide it the context to it. And see the results. For context, you are limited now by models contexts (1m), so mostly you would need to split what you have and prompt twice....or more...and so on
- inigyou 1mo agoSo uh ... Where's the retrieval part? You know RAG is used to implement that, right? You're basically saying "we don't need an ALU, we can just use the Windows calculator"
- respectattentio 1mo agoThe only difference is using LLMs instead of Embedding models
- seanspradlin0 1mo agoBut over-engineering things is fun. RAG is one of those things where I can hyper optimize to an absolutely needless degree.
- Wren_ops 1mo agoAgreed, simpler is almost always better. The hard part is resisting the urge to over-engineer it.
- eugenehizyyo 1mo ago[flagged]
- 13639366668 1mo ago[flagged]
- sonnykk19 1mo ago[flagged]
- saltysalt 1mo agoIf like me you run models locally, it's pretty easy to run your own RAG locally also using a Vector Database like Qdrant for persistence, and a middle-layer like Mem0 for realtime retrial and updates. I documented the set-up steps here: https://leadprompt.sh/a/739-Building-an-Infinite-Memory-Local-AI-Stack-on-Fedora https://leadprompt.sh/a/739-Building-an-Infinite-Memory-Loca...
- bityard 1mo agoThanks for the nice article. If you're looking for feedback, I'd suggest adding a short demo at the end. It would be nice to see you send it a prompt that says, "hey, remember this" and then tell it to recall that memory. Or show what the memories look like on their way to the model. Are the memories added to the context on every turn or only once per conversation?
- saltysalt 1mo agoThanks for the feedback, and that's a great idea I should have done that! They are extracted and added each turn, all handled by the same middleware proxy that also handles the retrievals. I put the full code for that (memory_proxy.py) in the article.
- Silasdev 1mo agoVery little of this is RAG but rather just FTS with clever reformulation and re-ranking. RAG is about providing an grounded response, given the actual data in the corpus. Great article and content, nonetheless!!
- j0selit0 1mo agoauthor here - thank you!
- Silasdev 1mo agoI apologize for write "just" FTS. I know there is a lot of work involved and your article summed it up pretty damn well, in a way that makes it approachable for someone new to the topic. I will keep a note of this article for next time I am asked about this topic.
- alansaber 1mo agoI have built systems using all of these approaches (all in tandem). For the most part, the juice is not worth the squeeze (in building a highly optimised corpus-specific information retrieval strategy) outside of a very few fringe cases. The amount of technical discussion far outstrips the use case for RAG.
- Tycho 1mo agoI don’t understand the 4th option, “on the fly”. It didn’t seem to be explained properly.
- j0selit0 1mo agoauthor here - apologies for it not being clear. the idea here is: step 1: sparse index retrieval (FTS/BM25) - say with k = 10 step 2: re-rank the 10 records using embeddings the difference in this approach is during step 2, you convert text to embeddings on the fly - when you're running the retrieval pipeline, meaning you don't need to have all of your corpus pre-embedded in a vector db
- klm127 1mo agoRAG stands for Retrieval Augmented Generation. The purpose is to search a corpus of text by meaning rather than exact match. I had to look it up.
- spunker540 1mo agoThat sounds more like semantic search and vector db. RAG is simply fetching external data (retrieval) and adding it to LLM context (augmenting) prior to generating a final response. Any time LLMs do a grep or a web search to answer the query, it’s RAG. Many people use vector db for their own RAG implementation bc of the semantic search benefits.
- 0x457 1mo agoBecause people writing about RAG never explained what RAG is and exclusively wrote about embeddings and vector dbs, for most people RAG became "embeddings + vector db". People don't understand that any sort of retrieval before generation is RAG.
- LowTechHN 1mo ago[dead]
- maxrumpf 1mo agoThe easiest way to strip complexity is to expose simple tools to an agent model like SID-1 that can use them well. It makes more of an effort for hard questions, and little effort for easy ones. (found of sid.ai so obv biased)
- hn58622tsf 1mo agoBookmarked, thanks again
- yipinwong 1mo agoOnly those who mastered the craft makes their work look simple. The AI that wrote this might be the master not the writer, as this looks written by AIs. I will use the author's agents, not read his articles or use him for the job.
- alankritxghoshx 1mo ago[flagged]
- waximabbax 1mo agoWe removed retrieval from our coding agent a while back. What convinced us wasn’t a benchmark, we found that the retrieval path had been returning zero results for quite some time because of a technical bug, still nobody noticed, indeed it was working better than before. After doing some rigorous A/B testing, we dropped indexing. For coding, I think the reason is that a repo is already searchable. Imports, call sites, file and test names, grep gives you cheap yet reliable version of what indexing would do, and the agent can read around a hit to verify it. Chunked retrieval hands the model something that looks right, and it tends to trust that instead of going to look for the actual source. Another thing that I noticed was the most intelligent models like Opus 5 and Fable ignored chunks anyway most of the time for some reason. Possibly perhaps they are trained around not trusting similarity checks for codebases. Extremely large codebases with docs feel different. You can’t grep for a concept you can’t name. That’s the case where I’d still use retrieval. (I work on TheGitAI, for disclosure.)
- entaroadun123 1mo ago[flagged]
- seamossfet 1mo agoI notice a lot of these AI written articles share this pattern where they'll present idea 1, then idea 2, and finally idea 3 which is some amalgamation of idea 1 and 2. Claude especially will present hybrid options and compromises to avoid having to make a choice then framing the hybrid option as the "best of both worlds" when they're borderline nonsensical. "on the fly embedding" and "Sparse + dense reranking" don't really make sense how they're presented and smell like they came from a long claude-driven conversation after multiple cycles of these hybrid compromises across many turns.
- ChipopLeMoral 1mo agoThis has Claude written all over it. "Recipe 4: On-The-Fly Embedding (The Fresh Data Play) The insight If your data changes frequently, why pay to re-embed everything?" This reads like every Claude generated presentation I've seen.
- seamossfet 1mo agoyeah, but I mean even prose specific claude-isms aside; the information itself is a weird patchwork of concepts
- Alifatisk 1mo agoI skimmed through the article and it seemed okay. But then I lost my enticement when reading the comments saying this is an LLM written article.
- geniium 1mo agoyet harder to implement proplery than you think
- akshay_akula 1mo agoAgreed. Embeddings are cheap to try and hard to mess up. Most projects can do plain semantic search first and see if they ever need more.
- _pdp_ 1mo agoAll computer primitives are relatively straightforward in pure form and vastly more complicated in real-world scenarios.
- zabriel_goss 1mo agoHelpful framing, thanks for sharing!
- maxweylandt 1mo ago> Elasticsearch. Postgres full-text search. The stuff that existed before “embedding” became a verb. Noun, no?
- garn810 1mo agoYet people sell vector DBs solutions as if it's a guarded magic knowledge Whole LLM agent tool call with ripgrep gives 99% use cases right lol
- ShinyLeftPad 1mo agoYou can just ask your LLM about it and get this article. This should've been a prompt.
- lhk931122 1mo ago[dead]
- nomad-linkd-id 27d ago[flagged]