2 ms·
This tutorial is a good overview of the rough systems behind most of this "Chat your data" application you're seeing now. https://www.pinecone.io/learn/langcha
by celestialcheese 3y ago
This tutorial is a good overview of the rough systems behind most of this "Chat your data" application you're seeing now.
https://www.pinecone.io/learn/langchain-retrieval-augmentation/ https://www.pinecone.io/learn/langchain-retrieval-augmentati...
Langchain / Llama indexes are both toolboxes that abstract away a lot of the plumbing for doing this kind of thing, and Pinecone is one of dozens of vector databases.
Personally, i'd try out langchain and chromadb and go through some of the examples langchain has in their docs, then be prepared to completely abandon langchain and just work with the LLM APIs directly. Start with openai, get on the waitlist for GPT-4 tokens, and also get on Anthropics Claude waitlist for the 100k-1.3. It's _very_ good for knowledge retrieval.
Langchain tries to do too much in extracting away the prompts, and the prompts are really what matter in getting interesting stuff out of your own data. Use langchain, llamaindex pieces but build from scratch for most things as your tinkering.
It's really not hard if you have a background in programming, and it's _so_ much fun. You'll feel like you have superpowers once you get a scraping interface hooked into an LLM. All of a sudden you can automate some really complex pipelines very quickly