14 ms·
LLMs, RAG, and the missing storage layer for AI
- saaaaaam 3y agoDoes ChatGPT always start articles with “in the rapidly evolving landscape of X”? Surely if you’re posting an article promoting miraculous AI tech you should human edit the article summary so that it’s not really obviously drafted by AI. Or just use the prompt “tone your writing down and please remember that you’re not writing for a high school student who is impressed by nonsensical hyperbole”. I’ve started using this prompt and it works astonishingly well in the fast evolving landscape of directionless content creation.
- jamesblonde 3y agoIt's not clear to me that only a vector DB should be used for RAG. Vector DBs give you stochastic responses. For customer chatbots, it seems that structured data - from an operational database or a feature store adds more value. If the user asks about an order they made or a product they have a question about, you use the user-id (when logged in) to retrieve all info about what the user bought recently - the LLM will figure out what the prompt is referring to. Reference: https://www.hopsworks.ai/dictionary/retrieval-augmented-llm https://www.hopsworks.ai/dictionary/retrieval-augmented-llm
- J_Shelby_J 3y agoAnd for technical documentation or code I'm unclear how well semantic search works for CEQ. I would assume the embedding model isn't trained on code and specific words that are industry/company specific.
- jarulraj 3y agoThanks for sharing that observation on customer chatbots. 1. Will that query look like this: SELECT LLM("{user_question}", order_info) FROM postgres_data.order_table WHERE user_id = “101”; 2. How will a feature store, like Hopsworks, help in this app? Shameless self-plug: We are building EvaDB [1], a query engine for shipping fast AI-powered apps with SQL. Would love to exchange notes on such apps if you're up for it! [1] https://github.com/georgia-tech-db/evadb https://github.com/georgia-tech-db/evadb
- jamesblonde 3y agoWhy would your projection be this - SELECT LLM("{user_question}", ? You can train a small llm on your private data to map the user question to tables in your db. Then Just select with a limit ( or time bounded). The feature store is just another operational store that could have relevant data for the query.
- fkyoureadthedoc 3y ago> You can train a small llm on your private data to map the user question to tables in your db. Can you? You've personally done this? Deployed it to production at some kind of non trivial scale and it's working well? I'm not aware of any "small llm" that approaches the quality of gpt-3.5.
- philipodonnell 3y agoThis is called Text2SQL or NL2SQL, it’s a surprisingly difficult problem even with RAG and GPT4 as soon as the query is non trivial, especially if there are semantic differences between the question and the db schema.
- Charon77 3y agoA lot of things mentioned are too handwaved and not explained well. It's not explained how vector DB is going to help while incumbents like chatgpt4 can already call functions and do API calls. It doesn't make AI less black box, it's irrelevant and not explained.. There's already existing ways to fine tune models without expensive hardwares such as using LoRA to inject small layers with customized training data, which trains in fractions of the time and resource needed to retrain the model
- antupis 3y agoThere is lots of things like which you don’t want leak eg customer specific data. For those cases vectors are great.
- panarky 3y agoThe first unstated assumption is that similar vectors are relevant documents, and for many use cases that's just not true. Cosine similarity != relevance. So if your pipeline pulls 2 or 4 or 12 document chunks into the LLM's context, and half or more of them aren't relevant, does this make the LLM's response more or less relevant? The second unstated assumption is that the vector index can accurately identify the top K vectors by cosine similarity, and that's not true either. If you retrieve the top K vectors according to the vector index (instead of computing all the pairwise similarities in advance), that set of 10 vectors will be missing documents that have a higher cosine similarity than that of the K'th vector retrieved. All of this means you'll need to retrieve a multiple of K vectors, figure out some way to re-rank them to exclude the irrelevant ones, and have your own ground truth to measure the index's precision and recall.
- NhanH 3y agoCould you please explain a bit on your 2nd paragraph. I couldn’t quite understand either the problem statement nor the reasoning itself.
- Jimmc414 3y agoSwitching to Word2Vec embeddings led to a substantial improvement in my cosine similarity evaluations for text similarity, but granted I was looking for actual similarity, not relevance. I tried many different methods and had lots of mediocre results initially. code: https://github.com/jimmc414/document_intelligence/blob/main/text_similarity.py https://github.com/jimmc414/document_intelligence/blob/main/... https://github.com/jimmc414/document_intelligence https://github.com/jimmc414/document_intelligence
- bugglebeetle 3y agoAs opposed to sentencebert or what?
- Jimmc414 3y agoDistilBERT and RoBERTa
- eth0pal 3y agoShameless self promotion
- juxtaposicion 3y agoWe use Lance extensively at my startup. This blog post (previously on HN) details nicely why: https://thedataquarry.com/posts/vector-db-4/ https://thedataquarry.com/posts/vector-db-4/ but essentially it’s because Lance is a “just a file” in the same way SQLite is a “just a file” which makes it embedded and serverless and straightforward to use locally or in a deployment.
- freedmand 3y agoI don’t fully understand the fascination with retrieval augmented generation. The retrieval part is already really good and computationally inexpensive — why not just pass the semantic search results to the user in a pleasant interface and allow them to synthesize their own response? Reading a generated paragraph that obscures the full sourcing seems like a practice that’s been popularized to justify using the shiny new tech, but is the generated part what users actually want? (Not to mention there is no bulletproof way to prevent hallucinations, lies, and prompt injection even with retrieval context.)
- deleted 3y ago[deleted]
- matchagaucho 3y agoIn a strict "one question / one response" search, raw semantic search results are a great solution. And consumes far fewer tokens. In conversational AI, providing search results appended to a long-memory context produces "human-like" results.
- sdenton4 3y agoOn the modeling side, it's compelling to separate the memory from the linguistic skills. Vector search is hella fast and can be very good. So you can off load the memorization part of the problem, and let the language model focus on the language. This should allow better performance with much smaller models.
- zawaideh 3y agoSometimes what I want is to ask Google/Alexa/Siri a question and get a summary response along with the source. I think that would be a good application of the above. Less so IMO when I’m on my phone or in front of the computer.
- nottheengineer 3y agoI really like using LLMs to learn stuff because they can explain anything at the exact level I need. Hallucination is a big problem with that and RAG pretty much solves it. If I give chatGPT a good stackoverflow post and tell it to dumb it down for me, it does very well. RAG just automates that process with the added benefit of not letting the LLM decide which information to retrieve, which should greatly reduce the chance of accidentally biasing the model with your prompt.
- dr_dshiv 3y ago404
- Binarybuilder 3y ago[flagged]
- ianpurton 3y agoAs an architect working on LLM applications I have these criteria for a database. - Full SQL support - Has good tooling around migrations (i.e. dbmate) - Good support for running in Kubernetes or in the cloud - Well understood by operations i.e. backups and scaling - Supports vectors and similarity search. - Well supported client libraries So basically Postgres and PgVector.
- ofermend 3y agoYes totally agree with that (and other comments below). Moving from a toy example to production deployment requires all the things we are used to having in robust/mature products like postgres.
- seanhunter 3y agoExactly. The whole point about databases is you don't need "a database for AI" you need a database, ideally with an extension to add additional AI functionality (ie postgres and pgvector). Trying to take a special store you invent for AI and retrofit all the desirable things you need to make it work properly in the context of a real application you're just going to end up with a mess. As a thought-experiment for people who don't understand why you need (for example) regular relational columns alongside vector storage, consider how you would implement RAG for a set of documents where not everyone has permission to view every document. In the pgvector case it's easy - I can add one or more label columns and then when I do my search query filter to only include labels that user has permission to view. Then my vector similarity results will definitely not include anything that violates my access control. Trivial with something like pgvector - basically impossible (afaics) with special-purpose vector stores. Or think about ranking. Say you want to do RAG over a space where you want to prioritise the most recent results, not just pure similarity. Or prioritise on a set of other features somehow (source credibility whatever). Easy to do if you have relational columns, no bueno if you just have a vector store. And that's not to mention the obvious things around ACID, availability, recovery, replication, etc.
- tinco 3y agoCan I add one more nice to have? Good support for graph data. I'm not 100% certain on it yet, but there's a lot of ideas surrounding storing knowledge as a graph out there and it makes a lot of intuitive sense. I haven't found a killer use case for it yet as so far I can get by just tagging things and sql querying on the tags is powerful enough. Maybe someone could pitch in. Is knowledge really a graph (for your problem domain), or is that just some bullshit people made up when they still thought AI could be captured mathematically? It feels to me now knowledge is much more like the way vector embeddings work, it's in a cloud where things are related to each other in an analog or statistical way, not a discrete way. But, perhaps for similar reasons, vector embeddings haven't been super useful to me in building RAG agents yet. Knowledge is either relevant or it's not, and at least for me if it's relevant it has the keywords or tags I need, and just a straight up SQL query brings it in.
- amelius 3y agoUnrelated question: is there a standard way for writing down neural network diagrams? I'm thinking of how it is done in electrical circuit schematics, which capture all relevant information in a single diagram, in a (mostly) standardized way. I've seen the diagrams in DL papers etc. but I guess everyone invents their own conventions, and the diagrams often don't convey the complete flow of information.
- gillesjacobs 3y agoThere are conventions and most libraries have libraries to export diagrams to LaTex or image (e.g., TorchViz). Visualizations are highly context and usage dependent anyway. Generally, there's is no value in showing fully connected or feed forward layers in detail outside of teaching materials.
- amelius 3y ago> Generally, there's is no value in showing fully connected or feed forward layers in detail outside of teaching materials. Well, in electrical circuit diagrams it is customary to draw e.g. a signal bus as a single connection, with the number of wires in the bus written next to it (with a little strike-through line). I'm guessing something similar can be done for DL networks.
- zwaps 3y agoI find it quite comical to speak of a "missing storage layer" during your own self-promotion, considering that the market for vector databases is literally overflowing right now. Everything else may be missing, but not the storage layer.