2 ms·
I recently had some luck turning an excel tracker that lists multiple locations and their services into markdown for RAG. It worked great as a natural language
by its_down_again 2y ago
I recently had some luck turning an excel tracker that lists multiple locations and their services into markdown for RAG. It worked great as a natural language lookup, way better than digging through a big Excel sheet.
I uploaded them through Supabase Embeddings Generator if you're curious.
https://github.com/supabase/embeddings-generator https://github.com/supabase/embeddings-generator
But things got a bit messy when I handed it off to someone else. They started using synonyms for locations, like abbreviated addresses to refer to certain columns, which didn't return the right documents.
Followed a friend's suggestion to try NotebookLM, so I uploaded the same docs there, and it was awesome. Some cloud-hosted vector DB tools only handle PDFs, but NotebookLM accepted my Markdown and chunked the docs better than the Supabase library I was using. It just "worked".
I would swap over to NotebookLM because their document chunking and RAG performance is working for my use case, but they just don’t offer an API yet.
I also gave Gemini a shot using this guide, but didn’t get the results I was hoping for.
https://codelabs.developers.google.com/multimodal-rag-gemini#0 https://codelabs.developers.google.com/multimodal-rag-gemini...
Am I overhyping NotebookLM? I’d love to know to get on-par document chunking, because that seems to deliver fantastic RAG right out of the box. I’m planning to try some other suggestions I’ve seen here, but any insights into how NotebookLM does its magic would be super helpful.