9 ms·
What is RAG? That's hard to search for
by seedless-sensat 3y ago
What is RAG? That's hard to search for
- deleted 3y ago[deleted]
- accrual 3y agoRetrieval-augmented generation, RAG + LLM will turn up more results.
- rjzzleep 3y agoThis one seems like a good summary Retrieval-Augmented Generation for Large Language Models: A Survey https://arxiv.org/abs/2312.10997 https://arxiv.org/abs/2312.10997 The photos of this post are also good for a high level look https://twitter.com/dotey/status/1738400607336120573/photo/2 https://twitter.com/dotey/status/1738400607336120573/photo/2 From the various posts I have seen people claim that phi-2 is a good model to start off from. If you just want to do embeddings, there are various tutorials to use pgvector for that.
- nmstoker 3y agoSeems fairly easy to search for to me - top results are all relevant: https://kagi.com/search?q=ml+rag https://kagi.com/search?q=ml+rag https://www.google.com/search?q=ml+rag https://www.google.com/search?q=ml+rag
- tomduncalf 3y agoRetrieval Augmented Generation - in brief, using some kind of search to find relevant documents to the user’s question (often vector DB search, which can search by “meaning”, by also other forms of more traditional search), then injecting those into the prompt to the LLM alongside the question, so it hopefully has facts to refer to (and its “generation” can be “augmented” by documents you’ve “retrieved”, I guess!)
- whartung 3y agoSo, as a contrived example, with RAG you make some queries, in some format, like “Who is Sauron?” And then start feeding in what books he’s mentioned in, paragraphs describing him from Tolkien books, things he has done. Then you start making more specific queries? How old is he, how tall is he, etc. And the game is you run a “questionnaire AI” that can look at a blob of text, and you ask it “what kind of questions might this paragraph answer”, and then turn around and feed those questions and text back into the system. Is that a 30,000 foot view really of how this works?
- Kubuxu 3y agoThe 3rd paragraph missed the mark but previous ones are in the right ballpark. You take the users question either embed it directly or augment it for embedding (you can for example use LLM to extract keywords form the question), query the vector db containing the data related to the question and then feed it all of LLM as: here is question form the user and here is some data that might be related to it.
- fennecbutt 3y agoEssentially you take any decent model trained on factual information regurgitation, or well any decently well rounded model, a llama 2 variant or something. Then you craft a prompt for the model along the lines of "you are a helpful assistant, you will provide an answer based on the provided information. If no information matches simply respond with 'I don't know that'". Then, you take all of your documents and divide them into meaningful chunks, ie by paragraph or something. Then you take these chunks and create embeddings for them. An embedding model is another type (not an llm) that generates vectors for strings of text often based on how similar the words are in _meaning_. Ie if I generate embeddings for the phrase "I have a dog" it might (simplified) be a vector like [0.1,0.2,0.3,0.4]. This vector can be seen as representing a point in a multidimensional space. What an embedding model does with the word meaning is something like if I want to search for "cat" that might embed as a vector [0.42]. Now, say we want to search for the query "which pets do I have" first we generate embeddings for this phrase, the word "pet" might be embedded as [0.41] in the vector. Because it's based on trained meaning, the vectors for "pet" and for "dog" will be close together in our multidimensional space. We can choose how strict we want to be with this search (basically a limit to how close the vectors need to be together in space to count as a match). Next step is to put this into a vector database, a db designed with vector search operations in mind. We store each chunk, the part of the file it's from and that chunks embedding vector in the database. Then, when the LLM is queried, say "which pets do I have?", we first generate embeddings for the query, then we use the embedding vector to query our database for things that match close enough in space to be relevant but loose enough that we get "connected" words. This gives us a bunch of our chunks ranked by how close that chunks vector is to our query vector in the multidimensional space. We can then take the n highest ranked chunks, concatenate their original text and prepend this to our original LLM query. The LLM then digests this information and responds in natural language. So the query sent to the LLM might be something like: "you are a helpful assistant, you will provide an answer based on the provided information. If no information matches simply respond with 'I don't know that' Information:I have a dog,my dog likes steak,my dog's name is Fenrir User query: which pets do I have?" All under "information" is passed in from the chunked text returned from the vector db. And the response from that LLM query would ofc be something like "You have a dog, its name is Fenrir and it likes steak."
- FergusArgyll 3y agoOff Topic; It fascinates me how much variance there is in peoples searching skills. some people think they are talking to a person when searching e.g 'what is the best way that i can {action}' I think the number one trick is to forget grammar and other language niceties and just enter concepts e.g. 'clean car best'
- szundi 3y agoThat’s why they will love chatgpt
- BlueGh0st 3y agoI used to do this. Then when Google's search results started declining in quality, I often found it better to search by what the average user would probably write.
- user32489318 3y agoand what would an average user write?
- amelius 3y agoAn entire question instead of a bunch of keywords.
- CommieBobDole 3y agoOver the last couple of years, at least with Google, I've found that no strategy really seems to work all that well - Google just 'interprets' my request and assumes that I'm searching for a similar thing that has a lot more answers than what I was actually searching for, and shows me the results for that.
- feitingen 3y agoSome concepts seems to be permanently defined as a spelling error and will just be impossible to search for.
- gianpaj 3y agoAsk chatgpt next time. "What is rag in context of AI?"
- pc86 3y agoOr just using a traditional search engine and "rag" plus literally any ML/AI/LLM term will yield a half dozen results at the top with "Retrieval-augmented generation" in the page title.
- rahimnathwani 3y agoOr if GGP can't think of an AI-related term they can use HN search. Searching 'rag' shows the term on the first page of results: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=Rag&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
- Sakos 3y agoSearching for "RAG" on Kagi and Google give some AI-related results fairly high up, including results that explain it and say what it stands for.
- valval 3y agoRight? How does someone who browses this forum not know how to find knowledge online?
- wackget 3y agoOr people could just not use obscure acronyms when discussing specialised topics on an open forum?
- robluxus 3y agoWhere do you draw the line though?
- gosub100 3y ago
- prestonlibby 3y ago"Retrieval augmented generation". I found success from "rag llm tutorial" as a search input to better explain the process.
- ksjskskskkk 3y agoRAG: having a LLM spew search queries for you because your search foo is worse than a chat bot alucinations. or because you want to charge your client the "ai fee". or because your indexing is so bad you hide it from your user and blame the llm assistant dept.
- namlem 3y agoI had the same query and instead of just scrolling down, I copy and pasted the paragraph into Bing chat and asked it what it meant. It got it right, but I probably should have scrolled farther first lol. It's retrieval augmented generation