4 ms·
Am I the only one who doesn't need to search across my data? What are the use cases here
by cloudking 3y ago
Am I the only one who doesn't need to search across my data? What are the use cases here
- BeetleB 3y agoExample use case: We have a group at work that meets and discusses various investment topics. The guy organizing it is fairly well connected and every week he tries to get an external speaker to come and present. Very educational. I have raw notes for each of these presentations. My goal has always been to go through those notes, and properly organize the knowledge in there into a wiki of sorts. It's been 3 years since this all started and I still haven't found the time to do it. If I want to be realistic, I should accept that it'll never happen. How do I go about finding information that I have in those notes? I could use text search but it's too sensitive to my search string - I'll often fail to find what I need. Also, the information may be scattered across several files, and I'd have to open all the hits and scan to find what I need. With technology like this, I can put all my notes into some vector DB, and then use AI to ask in plain English what I need. Locally the system interprets my query and finds the most relevant documents in the DB. It then sends my query and those hits to OpenAI to interpret my question, and find the answer amongst my notes. A while ago I used Langchain to set it up and I got it working as a proof of concept. An Aha moment was when I asked it something and it gave me a response with information that was scattered over two different presentations. My challenge is that there are so many parameters I could play with, and I haven't yet thought of a way/metric to assess the performance of my system (any pointers would be appreciated!) There's nothing personal in these notes, so no privacy concerns. I did want to set a similar thing up with over 20 years of emails, but didn't due to privacy. Also, I use a mail indexer (notmuch) which is fairly good so the need to use AI is not as strong. But for other (non-personal) notes? If I can get this system working fairly well, it'd be a life saver. I've made so many notes on so many topics over the years, and it's worth real money not to have to organize it well. Just let me write my notes, and use an AI to retrieve what I need.
- theonlybutlet 3y agoYou're creating additional hardship for yourself. Why create a pdf only to convert it out of pdf again. Just insert all your notes into the LLM model.
- muspimerol 3y agoBecause that requires retraining the model every time you take new notes. And this way you also still have the raw notes as similarity matches from the vector db, rather than them "disappearing" into the LLM model.
- theonlybutlet 3y agoI see thanks for the insight.
- BeetleB 3y agoYou are probably replying to the wrong thread - I didn't say anything about a pdf.
- theonlybutlet 3y agoApologies
- dkh 3y agoSometimes I have the data, but I'm not sure where it is. Sometimes I know where the data is, but there's a lot of it and all I'm looking for is a quick explanation of something. Sometimes I have a lot of data from a lot of sources, but what I want in the end is a summary based on what most/all of them agree on, or possibly a summary of how they differ. There's a lot of use-cases here, many of which I think people don't get a "lightbulb moment" about their usefulness until they've dug in and seen what is possible, because we are so used to how we approach these tasks normally. But the range of uses is quite broad. A project I'm working on for myself is a variation of this, where I've ingested years and years of my own notes and journals, and make queries for the purposes of my own introspection and personal growth. (I think there's a lot of of potential in this arena in general)