6 ms·
Show HN: I put PubMed in a vector DB
Hi HN,
As a researcher, I often found myself struggling with the limitations of keyword-based search when exploring PubMed papers. To address this, I created PubMed Search (https://www.pubmedisearch.com/ https://www.pubmedisearch.com/), a tool that leverages a vector database to enable semantic search across medical research literature.
Some key features:
* Daily updates to ensure access to the latest articles
* Semantic search using latest & greatest embedding models
* Some additional useful info about the papers (tldr, journal, publication date, etc.)
Hope you find it useful!
- rkwz 2y agoCongrats on shipping! I'm curious how the search results rankings work, doesn't look like it's based on date or number of citations, but seems to be deterministic (persists over multiple searches). I did a keyword search using one word.
- mpmisko 2y agoThanks! It uses a vector search approach. Your query is embedded in a vector space using a language model and we find the closest vector to the query from the PubMed papers. This is a good summary of the techniques: https://learn.microsoft.com/en-us/azure/search/vector-search-overview https://learn.microsoft.com/en-us/azure/search/vector-search.... There are a couple more tricks but this is the gist. The nice part is that this approach allows you to find relevant papers to your question. E.g, you can ask "Can secondhand smoke cause AMD?" and the very first few papers are answering your question (https://pubmedisearch.com/share/Can%20secondhand%20smoke%20cause%20AMD%3F https://pubmedisearch.com/share/Can%20secondhand%20smoke%20c...). The more specific question, the better. :)
- grumpopotamus 2y agoWhat are you embedding exactly? Chunks of documents?
- mpmisko 2y agoYes
- kkielhofner 2y agoNice! Out of curiosity what model(s) are you using to generate the embeddings?
- mpmisko 2y agoGlad you like it! I did this as a mini-project within our startup MediSearch (https://medisearch.io/ https://medisearch.io/) & the search pipeline is custom tuned for the problem.
- lucas_crocker 2y agoThis is very cool! 2 questions spring to mind: 1. How much did it cost to embed all those vectors and how many articles did you process? PMC is quite large. 2. Could elaborate a little more on your approach to ranking articles? Because I'm familiar with semantic search via embeddings put did you weight those with impact factors/citations? Like how does one even calculate that? Anyhow, love the idea.
- mpmisko 2y ago1. We cover all the articles on PMC. The exact cost is hard to estimate because we did a lot of iterations. 2. We do weight those ... it is a lot of trial and error and you have to have good & exhaustive benchmarks.
- bdangubic 2y agoHey mate, should search by PMID work? Like 35982160 is PMID for "Rare coding variation provides insight into the genetic architecture and phenotypic context of autism" - not seeing this publication at all in search results...
- mpmisko 2y agoHi, it currently does not support search by PMID. But you can find the paper included in the results here (5th place): https://pubmedisearch.com/share/Do%20some%20individuals%20with%20autism%20spectrum%20disorder%20(ASD)%20carry%20functional%20mutations%20rarely%20observed%20in%20the%20general%20population%3F https://pubmedisearch.com/share/Do%20some%20individuals%20wi...
- alex_duf 2y agoWhat storage did you go for, and what search approach?
- mpmisko 2y agoWe use pinecone and it is not ideal, looking at https://turbopuffer.com/ https://turbopuffer.com/ now. They look quite promising :)
- dpifke 2y agoVery cool! Related: the NIST TREC (Text REtrieval Conference) has had several competitions over the years related to improving the searchability of medical data: https://www.trec-cds.org/ https://www.trec-cds.org/ If you have novel ideas in this area, you should consider participating. https://trec.nist.gov/ https://trec.nist.gov/
- mpmisko 2y agoThanks! Looks quite relevant
- drycabinet 2y agoMaybe a stupid question, but how do you compare this against GPT-based search engines?
- mpmisko 2y agoGPT-based search engines usually use some sort of a database to retrieve context for the LLM to summarize first. This is what people refer to as RAG these days: https://blogs.nvidia.com/blog/what-is-retrieval-augmented-generation/ https://blogs.nvidia.com/blog/what-is-retrieval-augmented-ge.... Some of these GPT engines maintain their own vector DB to do semantic search, others are directly hooked into Bing / Google. So pubmedisearch.com would be one component of a GPT-based engine. We actually have a GPT-based engine here: https://medisearch.io/ https://medisearch.io/.
- madhatter999 2y agoVery promising tool based on a couple of questions I asked it! How did the cleaning of documents look like?
- mpmisko 2y agoLots of annoying edge cases as you can imagine, nothing particularly glamorous.
- mharig 2y agoAll 10 thumbs up! Edit: One suggestion: in the results list, please make the headings links to the articles, too.
- mpmisko 2y agoDone! Let me know if you have other feedback.