13 ms·
I want flexible queries, not RAG
- humanlity 2y agoif the data is all structured and can be easily trained, maybe the ML can be even better
- ashu1461 2y agoI think the author is over generalising things. Every tech is good for few things and not good for others. Vector Search, LLMs all are revolutionary technologies and have their own limitations but we should not form a bias on the basis of few edge cases.
- martin293 2y agoI do not understand the point of this article. You did not even tell us what the correct recipe was called. But let's ignore that for now. I did some googling around, and the Wikipedia article for sartu di riso [0] mentions the fact that "it (the dish) found success in Sicily". Also, in [1] a commenter going by the name of passalamoda mentions that they also make this dish in Sicily. The comment they wrote is in Italy, but Frank Fariello has translated it for us, or if you don't believe him for whatever reason, google translate does a fine job. All of this is to say that associating this dish with Sicily based on the short description you've given, is not far fetched at all. > I fail to see how an LLM summarizing the material would be an improvement. I am fairly confident that typing the question you gave ChatGPT, waiting a few seconds for an answer, and then reading it can easily take under a minute. Lets be lenient, and say it takes 5 minutes to also ask a second question and receive the answer. That would still take way way way less time then to find a book, get the book and go through the book to find the correct recipe. Also, you yourself have given a reason in your article as to why ChatGPT would be an improvement. I will quote it now: > Most of her time was dealing with the brittleness of the query interface (depending on word matching and source popularity), spam, and locked down sources. I have already spent way too much time on debunking a random internet article, but I also decided to try to get an answer from ChatGPT. I did that by continuing to ask it questions that a person looking for an answer, not contradictions, would ask. If we make the assumption, that deanputney is correct, and the dish you were looking for is Arancini al Burru, we are able to get an answer from ChatGPT by asking the very simple and natural question shown in [2]. [0] https://en.wikipedia.org/wiki/Sart%C3%B9_di_riso https://en.wikipedia.org/wiki/Sart%C3%B9_di_riso [1] https://memoriediangelina.com/2013/01/21/sartu-di-riso-neapolitan-rice-timbale/ https://memoriediangelina.com/2013/01/21/sartu-di-riso-neapo... [2] https://imgur.com/a/PKqGXyK https://imgur.com/a/PKqGXyK
- lukol 2y agoThanks for sharing your thoughts! I published a blog entry that describes our solution to this exact problem yesterday. We call it "Generation Augmented Retrieval (GAR)", please find the details here: https://blog.luk.sh/rag-vs-gar https://blog.luk.sh/rag-vs-gar
- slushy-chivalry 2y ago> To me, LLM query management and retrieval is much more valuable than response generation Sure, for certain tasks. For other tasks retrieval is less useful.
- mrbonner 2y agoHow do you delete or update entries in a vector database?
- Centigonal 2y agosame as in any other database
- ritzaco 2y agoI don't think the author completely understands RAG, and this article is a bit disconnected and unclear to me. Google already provides a "flexible natural language query interface". I think ironically it would be fairly trivial to build what he wants _using_ RAG. 1. Accept a natural language query, like ChatGPT et al already do 2. Ask an LLM to rephrase it in N different ways, optimized for Google searches 3. Scrape the top M pages for each of the output of 2 in parallel. You now have dozens of search results 4. Clean and vectorize all of these 5. Use either vector similarity or an LLM to return the best matching snippets for the original query from 1, constrained to stuff contained in 4. It would take a little longer than a ChatGPT or Google response, but I can see the appeal too.
- gkbrk 2y agoThis sounds exactly like what Perplexity.ai does.
- crote 2y agoHave you actually tried to use Google in the last few years? Its query interface is horrible! 2024-era Google doesn't give you what you ask for - it'll take a few keywords from your query and "helpfully" invent a completely different query. Try asking it for anything remotely obscure and it'll just return complete garbage. Unless you're looking for a "20 dishwashers under $300" article, Google is basically unusable these days. 10 or 15 years ago Google was actually useful. You could specify exactly what keywords to include - and more importantly exclude. What I am looking for (and probably the author too) is an LLM which rephrases a human-language request into a SQL-like query to feed to a 2010 Google. Bonus points if it actually clearly states the query it's going to run and ask for corrections.
- valicord 2y agoYou want google to search exactly for the keywords in the query without any "helpful" reinterpretation, but at the same time to accept flexible human language requests. How do you expect these 2 contradictory requirements to coexist?
- amath 2y agoThis is just that authors opinion. It seems clear that there is a spectrum of what users want from basic retrieval to content generation.
- regularfry 2y agoIs there something about the presentation of the article that gives you the impression this was intended to be read as other than the author's opinion?
- prionassembly 2y agoIs anyone doing something like using LLMs to generate Prolog (or Cyc, or some appropriately complex, brittle knowledge representation GOFAI)?
- surfingdino 2y ago> I wanted the retrieval of a good recipe, not an amalgam of “things that plausibly look like recipes.” And that's the core issue with AI. It is not meant to give you answers, but to construct output that looks like an answer. How is that useful I fail to understand.
- tedajax 2y agoDemos well, falls over in production after you've made the sale.
- throwawaymaths 2y agoSo LLMs are effectively a stack of (gradient descent) learned lookup/hash tables and you can build primitive but working logic systems out of lookup tables (think unidirectional turing machines) Can you elaborate on your claim "it is not meant to give you answers"?
- surfingdino 2y agoI am talking about what an ordinary person thinks an answer is. If the AI industry could pull its collective head of its ass it would notice that humans are looking for answers that are factually true and not probabilistically close to the concept of a plausible answer to the given question. To put it simply, if I ask my bank for a statement for the last month, I expect to see a list of actual transactions and I expect the numbers to add up, while AI fanboys are happy with output that looks like a bank statement, but has made up entries with random amounts paid in or out. I will likely get a lecture on how it's a wrong domain to apply AI to but the core objection stands--humans expect facts and order whilst AI cannot tell fact from made up stuff and will keep on generating fake answers that look like what humans are looking for. Humans are wired for survival and for processing information in a way that allows us to turn chaos into order, pattern, plan, or action script. This is why we find it so easy to spot AI generated content and why we reject it.
- throwawaymaths 2y agoOk. Thanks for your clarification, that makes sense. I see you're being downvoted, possibly because your tone is weird but I think it would do to know that he's not wrong... People expect an AI to both interpret fuzzy inputs and give definitive answers out. I call it the "Mr data fallacy". Unfortunately the reality is that garbage in, garbage out... Our inputs are garbage so there must be some likelihood of garbage. I think maybe the best we can hope for is to push garbage responses to the edges: when it gets it wrong it does it so obviously wrong (e.g. completely misformatted) that the answer is not acceptable to the user and obviously so.
- kunalgupta 2y agothis person just wants perplexity
- marcosdumay 2y agoI don't know if it's my queries, the situation of the web, or something internal to it, but my impression is perplexity is becoming less and less useful as time passes.
- more_corn 2y agolol the thing they’re asking for is literally rag. “I want a smart person to parse relevant information and return a concise and relevant answer” Take the query, find the book, dump the book into the context window and the LLM’s answer will be exactly what you want.
- deleted 2y ago[deleted]
- mpweiher 2y agoThis has been my thinking as well: the natural language interface is amazing and something we've been wanting for some time. The generation is a showy gimmick. So why aren't we separating the useful bit out? My sneaking suspicion is that we can't. It's a package deal in that there are no two parts, it's just one big soup of free-associating some text with other text, the stochastic parrot. LLM do not understand. They generate associated text that looks like an answer to the question. In order to separate the two parts, we'd need LLMs that understand. That, apparently, is a lot harder.
- marcosdumay 2y ago> So why aren't we separating the useful bit out? It was trained as a chat-bot so that's all the current training can do. If you want to use it, you must hack something useful out of the chat bot interface. It was trained as a chat bot because that was the impressive thing that got them investors. Useful applications need a lot more context to awe people, and context takes time and work to create. So, now that they've got money, are those companies that created those LLM chat-bots making a useful next generation engines behind the scenes? Well, absolutely not! The same situation applies on every researching round, and they need to show impressive results right now to keep having money. (And now I wonder... Why do VC investors exist again?)
- nomel 2y agoI think that's a short term perspective. > they need to show impressive results right now to keep having money Sure. You do as little as possible to make as much money as possible. This is a fundamental of commerce/human existence. But, at some point, it will end with everyone's models performing similarly, with free models catching up. The concept of sustained "impressive results" will eventually require actual "reasoning systems". They'll use all that accumulated wealth, from what you maybe perceive as low hanging fruit, to tackle it. I think it must be assumed that these AI companies are intentionally working toward that, especially since it's the stated goal of many of them. I think you must assume that these people are smart, and they can see the reality of their own systems.
- 2y ago
- jszymborski 2y agoIs that ChatGPT example representative of RAG? I thought ChatGPT was primarily generative. I think of something like Brave Search's AI feature when I think of RAG.
- jmount 2y ago(author) Good point, sorry about that. I guess I was think the generation is low value was a bit orthogonal to retrieval being very high value.
- danielmarkbruce 2y agoPeople are doing what you are are talking about. Use an LLM to parse non-structured data into a structured or semi-structured format (or even just create a bunch of structured metadata) that is more easily searchable, then using an LLM to parse a natural language query into a suitable query (maybe SQL), then find the result and return it.
- PaulHoule 2y ago… you really want to be able to cook documents down to facts, as in the old A.I., and then be able to make logical queries. Trouble is it is easy to ontologize some things (ingredients) but not so easy to ontologize the aspects of things that make things memorable.
- antirez 2y agoSuggested book about traditional Sicilian dishes: https://www.amazon.it/Profumi-Sicilia-libro-cucina-siciliana/dp/8886803737 https://www.amazon.it/Profumi-Sicilia-libro-cucina-siciliana...
- walterbell 2y agoWe need a trendy TLA for publisher and author marketing of books, highlighting their strengths relative to LLMs. HCL (Human Context Language) HLM (Human Language Model) LLH (Literate Language for Humans) LHM (Literate Human Model)
- jmount 2y agoThanks for the book recommendation! I found an English translation of a book by the author and I am ordering that.
- advisedwang 2y agoIsn't the point of RAGs to make (in this example) actual recipe databases accessible to the LLM? Wouldn't it get closer to the articles stated goal of getting the actual recipie?
- khaki54 2y agoYes. Normally with RAG you would actually search and try to pre-filter the data for the LLM.
- brigadier132 2y agoYes, I don't think the author fully thought through what they wrote. In essence they are saying they just want semantic search.
- cjf101 2y agoYes, but if you don't have the LLM at the end, a good search (against a good corpus with the needed info) would still have given the user what they wanted. Which in this case, is a human vetted piece of relevant information. The LLM really only would be useful in this case for dressing up the result and that would actually reduce the trust in the result overall. Alternatively a LLM could play a role as part of the Natural language pipeline that drives the search, hidden from the user, and I feel that that's a much more interesting use of them. The farther you go with RAGs, in my experience, the more they become an exercise in designing a good search engine, because garbage search results from the RAG stage always lead to garbage output from the LLM.
- creshal 2y ago> The farther you go with RAGs, in my experience, the more they become an exercise in designing a good search engine From what I've seen from internal corporate RAG efforts, that often seems to be the whole point of the exercise: Everyone has always wanted to break up knowledge silos and create a large, properly semantically searchable knowledge base with all intelligence a corporation has. Management doesn't understand what benefits that brings and doesn't want to break up tribal office politics, but they're encouraged to spend money on hypes by investors and golf buddies. So you tell management "hey we need to spend a little bit of time on a semantic knowledge base for RAG AI and btw this needs access to all silos to work", and make the actual LLM an after thought that the intern gets to play with.
- ilaksh 2y agoThe article would be more convincing if they showed the recipe from the book so we could compare it with the one that ChatGPT output. From a Google search, it looks like he's right about the poor accuracy. It gets the basic idea of the ingredients, but is not really accurate. And is initially wrong about the region. But actually, this is what RAG is for. You would typically do a vector search for something similar to the question about "rice baked in an egg mixture". And assuming it found a match on the real recipe or on a few similar possibilities, feed those into the prompt for the LLM to incorporate. So if you have a well indexed recipe database and large context window to include multiple possible matches, then RAG would probably work perfectly for this case.
- hn_throwaway_99 2y agoI'm also super curious to know what the actual recipe was. I literally just plugged the exact text from the article into ChatGPT 4o: > My mother remembers growing up with a Sicilian dish that was primarily rice baked in an egg mixture. Roughly a “rice frittata.” Do you know some examples of what dish or recipe this could be? And the response I got was > It sounds like your mother might be referring to a dish known as "Frittata di Riso" or "Frittata di Riso al Forno." This is a traditional Sicilian dish that combines leftover rice with eggs and other ingredients, then bakes it into a savory cake. Here is a basic recipe for Frittata di Riso: And then got a detailed, formatted recipe. I asked follow up questions along the lines of "Where did Frittata di Riso originate" and "Is Frittata di Riso popular in Sicily" and again got detailed, thorough answers. Fair enough that I don't know if it's hallucinating, but without more info from the author what can I compare it to?
- deanputney 2y agoI happen to have a copy of "Bruculinu, America" at home. Looking for rice recipes at roughly those positions in the book, it's either "Tumala d'Andrea: Rice Bombe" (blue bookmark) or "Arancini al Burru: Rice Balls with Butter". I'm going to guess it's Arancini al Burru, because the Tumala d'Andrea is way more involved (and stuffed with pasta)[0]. Here's a similar recipe for Arancini al Burru[1]. However, Arancini al Burru is fried and Tumala d'Andrea is baked. So I'm still speculating, it could be either or something different. It's a nice cookbook worth adding to your collection. [0] https://inasmallkitchen.wordpress.com/2012/08/12/tumala-dandrea/ https://inasmallkitchen.wordpress.com/2012/08/12/tumala-dand... [1] https://www.196flavors.com/arancini-al-burro/ https://www.196flavors.com/arancini-al-burro/
- WhatIsDukkha 2y agoYour input was incredibly low effort and prompted a very low effort output. I took part of your blog post (which you clear were willing to put a few more tokens into) - "My mother remembers growing up with a sicilian dish that was primarily rice baked in an egg mixture. Roughly a "rice frittata". What are some distinctly Sicilian dishes that this could be referring to?" Notice there is not much extra context that you've offered any of us, either the LLM or us. You didn't even tell us what the recipe was... How was the dish served, what did it look like? What are you expecting of the LLM here? It not a psychic AGI.
- jmount 2y agoI guess I should have emphasized how I didn't like the LLM contradicting itself.
- WhatIsDukkha 2y agoI would expect most any answer to be loopy and partially wrong (or really I'm not surprised when they are) based on the short question, its just something you build intuition around as you use them. edit - btw I just noticed "Sartù di Riso: A baked rice dish that can include ingredients like meat, peas, and cheese, often bound together with eggs. It’s more commonly associated with Naples but has variations in Sicily." Was one of the 4 dishes in the question I submitted. So... was it contradictory actually?
- Dylan16807 2y ago> I would expect most any answer to be loopy and partially wrong (or really I'm not surprised when they are) based on the short question, its just something you build intuition around as you use them. That sounds pretty unpleasant as a rule to follow. How should I ask instead? Can I make an LLM do the question-expansion for me?
- LASR 2y agoI continue to hold the strong position that calling LLMs without injecting source truths is pointless. LLMs are exceptionally powerful as a reasoning engine. It’s useless as a source of truths or facts. We have chat bots, chat bots with automatic RAG etc. After the initial excitement wears off, you’re going to want a way to inspect and adjust the source queries yourself. In this case, being able to select what to search for in Google might be a good way for the cooking recipe usecase.
- crabmusket 2y agoFor those who haven't read it yet: "I like to think of language models like ChatGPT as a calculator for words. This is reflected in their name: a “language model” implies that they are tools for working with language. That’s what they’ve been trained to do, and it’s language manipulation where they truly excel. Want them to work with specific facts? Paste those into the language model as part of your original prompt!" https://simonwillison.net/2023/Apr/2/calculator-for-words/ https://simonwillison.net/2023/Apr/2/calculator-for-words/
- schmidt_fifty 2y ago[dead]
- pseudosavant 2y agoI think complaints like this show just how amazing AI is getting. This person really expected that ChatGPT would single-shot give them this obscure recipe that took them a ton of effort to find themselves. Current AI can do so much, that people lament that it can't do everything. It is incredible to me when people bring up bad AI generated legal fillings... like people actually expect it to already to all of the work of an attorney, without error.
- deleted 2y ago[deleted]
- Repulsion9513 2y agoIt's not that it can't do everything. It's that it pretends it can do everything.
- stavros 2y agoIt's like that old joke about one person being amazed at someone's dog being able to sing, and the other person says "it's not that impressive, he's pitchy" or something. Computers can finally think, and we're annoyed they can't think better than us.
- RevEng 2y agoI think of it more as "dancing bear ware". If you've ever seen a dancing bear, you wouldn't say it's particularly good at dancing. You may even say it's hardly dancing at all. But we don't care that it dances well; it's amazing that it dances at all. Current AI is a dancing bear. It gives us the feeling of understanding language, semantics, logic and reason, but when you look closely, you realize it's doing a very poor job of it in a way that suggests it is just mimicking those things without actually being capable of them.
- pseudosavant 2y agoExcept for a dancing bear is useless. I get a ton of real world use out of GPT-4 that no other kind of product in this world can do.
- bambax 2y ago> There is a lot of excitement around retrieval augmented generation or “RAG.” Roughly the idea is: some of the deficiencies in current generative AI or large language models (LLMs) can be papered over by augmenting their hallucinations with links, references, and extracts from definitive source documents. I.e.: knocking the LLM back into the lane. This seems like a misunderstanding of what RAG is. RAG is not used to try to anchor to reality a general LLM by somehow making it come up with sources and links. RAG is a technology to augment search engines with vector search and, yes, a natural language interface. This concerns, typically, "small' search engines indexing a specific corpus. It lets them retrieve documents or document fragments that do not contain the terms in the search query, but that are conceptually similar (according to the encoder used). RAG isn't a cure for ChatGPT's hallucinations, at all. It's a tool to improve and go past inverted indexes.
- hn_throwaway_99 2y agoThis feels like a "No true Scotsman" answer. You say > RAG is not used to try to anchor to reality a general LLM by somehow making it come up with sources and links. but that definitely is one particular use of RAG, i.e. to limit some potential hallucinations by grounding it in data provided in the prompt.
- bambax 2y agoNot really. RAG works very differently from a general (or generalist?) LLM. RAG is vector search first. It encodes the query, finds nearest vectors in the vector database, retrieves the fragments attached to those vectors, and then sends those vectors to the LLM for it to summarize them. A general LLM like Gemini or Claude or ChatGPT first produces an answer to a question, based on its training. This doesn't involve searching any external source at that point. Then after that answer is produced, the LLM can try to find sources that match what it has come up with.
- outofpaper 2y agoThis is a generalization. These proprietary systems do different things at different times. With GPT4o you can see little icons when a RAG is in use or when code and tests are being used. People, we have to stop talking about what we know as though it's all there is. Don't confuse our knowledge for understanding. Understanding only comes from repeadly trying to prove our understandings wrong and learning how things truly are.
- khaki54 2y agoClearly this author doesn't know what RAG is. RAG would be if he first did a 'retrieval' of all his mother's cookbooks containing Italian recipes, then reviewed the index for rice, scanned and OCRd those pages. That data would be submitted to ChatGPT with the query to 'augment' it within the constraints of the context window so that ChatGPT could generate a response with the highly relevant cookbook info.
- budududuroiu 2y agoAn interface for vector search is 100x more helpful to me than an LLM spitting out the same content as slop. The key to vector search is how you chunk your data, but I have some libraries to help with that
- lmeyerov 2y agoTo keep control in the hands of the analyst, we've been working on UX's over agentic neurosymbolic RAG in louie.ai -- Ex: "search for login alerts from the morning, and if none, expand to the full day" That requires generating a one-shot query combining semantic search + symbolic filters, and an LLM-reasoned agentic loop recovering if it turns up not enough such as a poorly formed query around 'login alerts' and the user's trigger around 'if none' Likewise, unlike Disneyified consumer tools like chatgpt and perplexity that are designed to hide what is happening, we work with analysts who need visibility and control. That means designing search so subqueries and decisions flow back to the user in an understandable way: they need to inspect what is happening and be confident they missed nothing, and edit via natural language or their own queries when they want to proceed Crazy days!
- civilized 2y agoVery good to hear this. As a data analyst, I have tended to dismiss LLMs as irrelevant because of the black-box mentality. In some industries, this even holds back adoption of technology that is now considered mature and boring, like tree ensembles in machine learning.
- radarsat1 2y agoOnce tree ensembles get big enough to handle the kinds of problems that LLMs can address, are they really more interpretable?
- thom 2y agoYup, this is the way. Natural language is an excellent search/specification front end. But without disambiguation and perfect clarity on how a query was interpreted, you cannot trust a black box for real work.
- jillesvangurp 2y agoThe retrieval part of RAG is usually vector search; but it doesn't have to be. Or at least not exclusively. I've worked with various search backends for about 20 years. People treat vector search like magic pixie dust but the reality is that it's not that great unless you heavily tune your models to your use cases. A well tuned manually crafted query goes a long way. Pretty much any system I've built over the last few years, the best way to think about search is about building a search context that includes anything relevant to answering the user's question. The user's direct input is only a small part of that. In the case of mobile systems, the user entered query is actually typically a very minor part of it. People type two or three letters and then expect magic to happen. Vector search is completely useless in situations like that. Why does search on mobile work anyway? Because of everything else we know to create a query (user location, time zone, locale, past searches, preferences, etc.) RAG isn't any different. It's just search where the search results are post processed by an LLM with whatever the user typed. The better the query and retrieval, the better the result. The LLM can't rescue a poorly tuned search. But it can dig through a massive result of search results and extract key points.
- cnity 2y agoThe valuable part _is_ the vector search, at least from the perspective of the OP. It sounds like what they're after is to just pass back the source texts from the vector search, rather than passing them back through an LLM to be summarised. I generally agree with the OP: in a search context I often don't want that last LLM summarisation step. I want embeddings generated from my search query for document lookup, and I want to see the original sources. Phind did (does?) something like this with citations, which is at least an improvement.
- danielbln 2y agoIt's common practice to have the LLM spit out the meta data as well, e.g. a source link. We do that for almost all of our usecases, it's the best way to increase trust in the response. The LLM step does provide a benefit, as it can synthesize a tight and concise answer to the user's question, instead of returning a list of source documents to sift through manually.
- valstu 2y agoWe use the term "pre-googling" for this sort of "information retrieval". You might have some concept in your head and you want to know the exact term for it, once you get the term you're looking for from LLM you'll move to Google and search the "facts". This might be a weird example for native english speakers but recently I just couldn't remember the term for graph where you're allowed to move in one direction and cannot do loops. LLM gave me the answer (directed acyclic graph or DAG)right away. Once I got the term I was looking for I moved on to Google search. Same "pre-googling" works if you don't know if some concept exits.
- jmount 2y agoThe pre-Googling is an excellent idea. You are augmenting the query, not generating nonsense answers. My wife uses ChatGPT as a thesaurus quite a lot.
- cj 2y ago> graph where you're allowed to move in one direction and cannot do loop To be fair, you didn't need LLM for this. Googling that, the answer (DAG) is in the title of the first Google result. (Not to invalidate your point, but the example has to be more obscure than that for this strategy to be useful)
- bigfudge 2y agoI recently started watching fallout and it reminded me of a book I read about a future religious order which was piecing together pre-bomb scientific knowledge. It immediately pointed me to the Canticles of Leibovitz (which is great btw). Google results will do the same, but llm I’d much faster and more direct. I find it great for stuff like this - where you know there is an answer and will recognise it as soon as you see it. I genuinely think it can become an extension of my long-term memory, but I’m slightly nervous about the effect it will have on my actual non-memory if I just don’t need to remember stuff like this anymore!
- immibis 2y agoUsers don't want what they think they want. Therefore, we won't give users what they think they want.
- a_c 2y agoQuerying knowledge is not a nail, but the hammer is generation. It is written on the tin, "generation" AI. People want "insight", "summary", "workflow automation", "code completion" from a guided proverbial monkey hammering on the keyboard hoping that our problem will become a nail. It is getting closer though