12 ms·
I used GPT to build a search tool for my second brain note-taking system
- PaulHoule 4y agoWould be nice to see some indication of how well it works in his case. I worked on a ‘Semantic Search’ product almost 10 years ago that used a neural network to do dimensional reduction and had inputs to the scoring function from the ‘gist vector’ and the residual word vector which was possible to calculate in that case because the gist vector was derived from the word vector and the transform was reversible. I’ve seen papers in the literature which come to the same conclusion about what it takes to get good similarity results w/ older models as a significant amount of the meaning in text is in pointy words that might not be included in the gist vector, maybe you do better with an LLM since the vocabulary is huge.
- sandkoan 4y agoI'd honestly argue that he might not have even needed OpenAI embeddings—any off-the-shelf Huggingface model would've sufficed. Because of attention mechanisms, we no longer so heavily depend on the existence of those "pointy words," so generally, Transformers-based semantic search works quite well.
- PaulHoule 4y agoI was thinking RoBERTa 3, longformer or Big Bird would be a good choice for this, though having any limit on the attention window is a weakness.
- lpasselin 4y agoI actually tried this last year, before OpenAI released their cheaper embeddings v2 in december. From my experiments, when compared to Bert embeddings (or recent variation of the model) the OpenAI embeddings are miles ahead when doing similarity search.
- leobg 4y agoInteresting. Nils Reimers (SBERT guy) wrote on Medium that he found them to perform worse than SOTA models. Though that was, I believe, before December.
- PaulHoule 4y agoMost of the practitioners I see attempting this are running text through an embedding and then using cosine similarity or something similar as a metric. Nils has written a lot of papers https://www.nils-reimers.de/ https://www.nils-reimers.de/ and I think the Medium post you are talking about is https://medium.com/@nils_reimers/openai-gpt-3-text-embeddings-really-a-new-state-of-the-art-in-dense-text-embeddings-6571fe3ec9d9 https://medium.com/@nils_reimers/openai-gpt-3-text-embedding... and that SBERT is a Siamese network over BERT embeddings https://arxiv.org/abs/1908.10084 https://arxiv.org/abs/1908.10084 which one would expect to do better than cosine similarity if it was trained correctly. I'd imagine the same Siamese network approach he is using would work better than cosine similarity with GPT-3. There's also the issue of what similarity means for people. I worked on a search engine for patents where the similarity function we wanted was "Document B describes prior art relevant to Patent Application A". Today I am experimenting with a content based recommendation system and face the problem that one news event could spawn 10 stories that appear in my RSS feeds and I'd really like a clustering system that groups these together reliably without false positives. I'd imagine a system that is great for one of these tasks might be mediocre for the other, in particular I am interested in some kind of data to evaluate success at news clustering.
- michaericalribo 4y agoThese "augmented intelligence" applications are so exciting to me. I'm not as interested in autonomous artificial intelligence. Computers are tools to make my life easier, not meant to lead their own lives! There's a big up-front cost of building a notes database for this application, but it illustrates the point nicely: encode a bunch of data ("memories"), and use an AI like GPT to retrieve information ("remembering"). It's not a fundamentally different process from what we do already, but it replaces the need for me to spend time on an automatable task. I'm excited to see what humans spend our time doing once we've offloaded the boring dirty work to AIs.
- pessimist 4y agoIn chess we had a tiny window of a few years when humans could use the help of computers to play the world's best chess. By 2000, computers were far better than humans and the gap has increased. Chess players are now entertainers, like all us humans are destined to spend our time doing.
- rlayton2 4y agoWhile reductive, isn't that true in many professional sports though? I have a wide variety of tools I can use to travel 100m faster than Usain Bolt, but its incredible to watch him do it on his own.
- l33t233372 4y agoI agree that professional athletes are entertainers, but I’m not sure what you’re getting at with that point.
- haswell 4y agoThe way I read this, in a world with machines that can travel at high speeds, people still watch professional runners because they are interesting to watch, and we’re inspired by human achievement. For similar reasons, it doesn’t really matter that computers are better at chess than us.
- sowbug 4y agoI wonder whether your individually trained chat bot will be allowed to assert the Fifth Amendment right against self-incrimination to stop it from talking when the police interview it. And if it's allowed, do you or it decide whether whether to assert it? What if the two of you disagree? Similar questions for civil trials, divorce proceedings, child custody....
- michaericalribo 4y agoImagine a model that decides on its own to assert the Fifth on your behalf. But now imagine that AI decides to lock you out of your own system...
- DavidPiper 4y agoWe already have those. They're called Google, Microsoft and Apple.
- bongoman37 4y ago[dead]
- qwertox 4y agoThis is a topic which really deserves a lot more attention, as in: from magazines to newspapers to talk shows. Seems like an appropriate time to get it on the agenda before governments opt to decide on their own.
- roywiggins 4y agoHow could it have a right not to self-incriminate when it can't be tried for a crime? An AI can't be indicted or convicted. Humans can be required to testify too if they're immunized.
- sokoloff 4y agoWith limitations. I cannot be compelled to testify against my wife (and probably not against my kids, though I’m unsure of that [edit: that seems to vary by state currently]), even if I personally am granted immunity.
- tra3 4y agoThis is fascinating. Can I train it on 5 years of stream of consciousness morning brain dumps and then say "write blah as me"? Before I do that, I'd love to know if training data becomes part of the global knowledge base available to everyone..
- michaericalribo 4y agoThese privacy considerations are highest-priority for any extended roll-out of LLM-based products. Privacy on the side of model servers would be good. Open source models that can be run locally would be better.
- feanaro 4y agoI personally think anything server-side is unacceptable. Only open source and local will fly.
- ilaksh 4y agoThis is not a fine-tuning example. It's an embedding search example. You use the embeddings to search for relevant knowledgebase chunks and then include them in the prompt. Which goes to the original model, not a model that you have trained more. This is popular because it's much much easier to do effectively than fine tuning and the OpenAI model is very capable of integrating kb snippets into a response. What I have heard is that it's easy to overdo fine tuning with OpenAI's model and makes more sense when you want a different format of response rather than just pulling in some content. Having said all of that, they do have a fine-tuning endpoint and I am guessing if you find the right parameters and give it a lot of properly formatted training data then it will be able to do an okay job. I have the impression it is not easy to do either of those things quite right though. As far as privacy, no they will not share your data when you use the API. ChatGPT is different, they ARE using the inputs to train the model.
- crosen99 4y ago> Having said all of that, they do have a fine-tuning endpoint and I am guessing if you find the right parameters and give it a lot of properly formatted training data then it will be able to do an okay job. Unfortunately, the fine-tuning API cannot be used to add knowledge to the model. It only helps condition the model to a certain response pattern using the knowledge it already has.
- leobg 4y agoSlight overkill to use GPT, though it works for the author and I can see that it’s the low hanging fruit, being available as an API. But this can also be done locally, using SBERT, or even (faster, though less powerful) fastText. Also, it’s helpful not to cut paragraphs into separate pieces, but rather to use a sliding window approach, where each paragraph retains the context of what came before, and/or the breadcrumbs of its parent headlines.
- dchuk 4y agoWhen using SBERT instead of gpt for this use case, is it paired with some sort of vector database or just all done in code/memory?
- leobg 4y agoYou’d want persistence, since the embedding process takes some time. But you don’t need to go all Pinecone on this. There is FAISS, and there is hnswlib, for example. Like SQLite for vector search.
- gk1 4y agoFriendly reminder that we (Pinecone) have a free tier that holds up to ~5M SBERT embeddings (x768 dimensions). For quick projects, going "all Pinecone on this" could turn out to be the easier and faster option.
- leobg 4y agoPoint taken ;-) I like to stand up for the little guy. I hear Pinecone this and Pinecone that. And nobody seems to pay any attention to the awesome dude who made hnswlib.
- gk1 4y agoWho, Yury Malkov? He won’t be offended… He’s an advisor to Pinecone. :) And yes, both he and HNSW are awesome.
- lukemtx 4y agoI wanted to do this! :D
- trane_project 4y agoI've been thinking of using GPT or similar LLMs to extract flashcards to use with my spaced repetition project (https://github.com/trane-project/trane/ https://github.com/trane-project/trane/). As in you give it a book and it creates the flashcards for you and the dependencies between the lessons. I played around with chatgpt and it worked pretty well. I have a lot of other things in my plate to get around first (including starting a math curriculum) but it's definitely an exciting direction. I think LLMs and AI are not anywhere near actual intelligence (chatgpt can spout a lot of good sounding nonsense ATM), but the semantic analysis they can do is by itself very useful.
- throwaway675309 4y agoI've seen a number of projects around using GPT to generate curriculum and also flashcards in the past three months, I think this is one of the most popular ones: https://autolearnify.com https://autolearnify.com It's a very good idea in theory but takes almost as much work to verify that the flashcards and curriculum that it generates is accurate and not a hallucinogenic nightmare. The biggest danger is that the target audience are not experts in the desired subject domain, so they have no way of sanity checking the generated curriculum.
- trane_project 4y agoWhen I played with it, I made it output a JSON file so that it would be easier to handle the output. And I specifically gave it the text to use. It did a pretty good job, but I ran into output size limits. I agree that using the training data would probably generate more garbage. But it's the semantic analysis part that I think it's useful. In general, I think VCs and OpenAI are overhyping it by calling it "intelligent" and obscuring the very good use cases of the technology. AFAIK, no one involved has explained how a statistical model running on a Turing machine magically develops agency and awareness, which are requirements for actual intelligence (under my definition, at least).
- james-revisoai 4y agoCurious as to where you came to this impression? I'd say the most popular applications are Knowt in the US for now, and Saveall.ai + Revision.ai (my company) in the UK, all been around with BERT/T5 etc long before this GPT trend. The flashcard accuracy varies wildly amongst current solutions, that's for sure.
- rolenthedeep 4y agoOne of my biggest dreams is a self-hosted AI that always listens through my phone and automatically takes notes, puts events in my calendar, set reminders, and template journal entries. A true personal assistant to keep my increasingly-complex life in order. I'd love a system where I can just point a search engine at my brain. I tried really hard for a while, but I just didn't have the discipline or memory to exhaustively document everything. An AI that can do this kind of thing in the background would be an absolute godsend for ADHD and ASD people.
- senectus1 4y agoMS is very close to this. they have chatGPT listening to meetings and taking notes then adding tasks to attendee calendars...
- lazyasciiart 4y agoWould be useful if it can insert notes like “no agenda planned” and “this meeting could have been an email”
- alostpuppy 4y agoLink? Because I’m gonna need this.
- senectus1 4y agothey lightly cover it in this blogpost https://www.microsoft.com/en-us/microsoft-365/blog/2023/02/01/microsoft-teams-premium-cut-costs-and-add-ai-powered-productivity/ https://www.microsoft.com/en-us/microsoft-365/blog/2023/02/0... but I've seen more detailed capability... I can't remember if It was under NDA though. I cant seem to find it with a quick search though
- tactiq 4y agohttps://tactiq.io/ https://tactiq.io/
- seeraan 4y ago
- abrkn 4y agoI’d love to have a ChatGPT that was also trained on all of the pages from my “second brain,” Roam Research. Imagine, I could ask it questions about myself, my friends, and my business. It would in many ways know me better than me from reading all my journal entries. How many years are we away from something like this?
- DanielVZ 4y agoZero. I’ve done this with google drive and GPT-3 (thus quite limited in prompt length). The biggest hurdle for your requirements is the nonexistent public API for ChatGPT and Roam Research.
- kanyethegreat 4y agoObsidian is just plain text. Only a matter of time before ChatGPT is accessible via API (there are some libraries that have reverse engineered the API; I put ChatGPT in VSCode when it first came out). Until then, GPT3 has an API.
- mtnygard 4y agoNot sure if this is still supported, but Roam used to have an in-app query interface. You could use the JS console to run Datomic style Datalog queries.
- johntash 4y agoI was tempted to do something like this but couldn't get over the idea of sending a bunch of personal data off to openai (or any 3rd party really)
- deleted 4y ago[deleted]
- totetsu 4y agoI think letting a language model make an outline of a topic you want to make notes on and writing in the details might not be such a bad thing.
- 110 4y agoIn case folks are interested in trying it out, I just released the Obsidian plugin[1] for Khoj (https://github.com/debanjum/khoj#readme https://github.com/debanjum/khoj#readme) last week. It creates a natural language search assistant for your second brain. Search is incremental and fast. You notes stay local to your machine. There's also a (beta) chat API that allows you to chat with your notes[2]. But that uses GPT, so notes are shared with OpenAI if you decide to try that. It is not ready for prime time yet but maybe something to check out for folks who are willing to be beta testers. See the announcement on reddit for more details[3] Edit: Forgot to add that khoj works with Emacs, Org-mode as well[4] [1]: https://obsidian.md/plugins?id=khoj https://obsidian.md/plugins?id=khoj [2]: https://github.com/debanjum/khoj#chat-with-notes https://github.com/debanjum/khoj#chat-with-notes [3]: https://www.reddit.com/r/ObsidianMD/comments/10thrpl/khoj_an_ai_search_assistant_for_your_second_brain/?utm_source=share&utm_medium=web2x&context=3 https://www.reddit.com/r/ObsidianMD/comments/10thrpl/khoj_an... [4]: https://github.com/debanjum/khoj/tree/master/src/interface/emacs#readme https://github.com/debanjum/khoj/tree/master/src/interface/e...
- bostonvaulter2 4y agoThis looks great! I definitely plan on checking it out.
- whatever1 4y agoSo if you keep adding notes furiously every day for years, do you asymptotically get your consciousness on—a-chip?
- college_physics 4y agoDesktop search feels like it has stagnated for at least a decade. Yet its an obvious way to both enhance privacy, improve relevance and even open up entirely new capabilities
- articsputnik 4y agoWow, this was super interesting as someone using a Second Brain daily. Thank you so much for digging into it, putting in the work, and sharing with us all! Much appreciated. I will follow you for more. I am much excited to do more with my Second Brain, but one concern, as you point out, is to use chatGPT or similar; we'd need to upload all our private and sometimes sensitive notes, which is a no go for me. So happy that you do everything locally. I wonder what the equivalent would be to train the model to search and ask questions based on our second brain (plus the already trained information). That's also where Obisidan will win in the long run, as other tools do not have the data locally. Obviously, it's already in the cloud; they could train on them, but training on customer-sensitive data would be a big problem. Something I will follow closely.
- FiberBundle 4y agoDoes anybody know how search engines apply semantic search with embeddings? To my knowledge no practical algorithms exist that find nearest neighbors in high dimensional space (such as that in which word/sentence/document vectors are embedded in), so those wouldn't give you any benefit compared to an iterative similarity search as applied here. Which obviously is totally impractical for real search engines. There are approximate nearest neighbor algorithms such as Locality-sensitive hashing, but even they seem impractical for real world usage on the scale of the indexes that search engines use. So how can Google e.g. make this work?
- gk1 4y agoThere actually are practical algorithms for finding approximate nearest neighbors (ANN) at large scales. Some of them are open source like HNSW [1] and Faiss [2], and some are even offered inside managed services like Pinecone.[3] [1] https://www.pinecone.io/learn/hnsw/ https://www.pinecone.io/learn/hnsw/ [2] https://www.pinecone.io/learn/faiss/ https://www.pinecone.io/learn/faiss/ [3] https://www.pinecone.io/ https://www.pinecone.io/
- qwerty456127 4y agoI don't imply any judgment (like good/bad) but I tend to suspect the major (not necessarily intentional) reason of all these "second brains" (I use too) to exist in the grand scheme of things is to be a high-quality input for AIs to learn.
- Terretta 4y agoObsidian is offline static markdown files If offline static markdown exist to be input for AIs to learn, then OK. But for me, long before Obsidian, they're just notes, because the act of writing them causes you to process them "outbound" which is the same process you need to recall them later.
- danwee 4y agoUmm, the only thing that stops me from doing this is uploading my notes to OpenAIs' servers.
- asdff 4y agoExactly. They should be paying you for the training data you've given them not the other way around.
- asdff 4y agoDid the author show how this system outputted results? I see an example of a lexical search and the technical implementation, but no example of some semantic output showing how its relevant to the lexical search string without containing that string. The author used the literal search string "failure mode" as their example. I was wondering if chatgpt would bring up results relevant to the lay person interpretation of failure mode, a technical interpretation, or something in between.
- gokulkrishh09 4y ago[dead]