4 ms·
Show HN: Evidex – AI Clinical Search (RAG over PubMed/OpenAlex and SOAP Notes)
Hi HN,
I’m a solo dev building a clinical search engine to help my wife (a resident physician) and her colleagues.
The Problem: Current tools (UpToDate/OpenEvidence) are expensive, slow, or increasingly heavy with pharma ads.
The Solution: I built Evidex to be a clean, privacy-first alternative. Search Demo (GIF): https://imgur.com/a/zoUvINt https://imgur.com/a/zoUvINt
Technical Architecture (Search-Based RAG): Instead of using a traditional pre-indexed vector database (like Pinecone) which can serve stale data, I implemented a Real-time RAG pattern:
Orchestrator: A Node.js backend performs "Smart Routing" (regex/keyword analysis) on the query to decide which external APIs to hit (PubMed, Europe PMC, OpenAlex, or ClinicalTrials.gov).
Retrieval: It executes parallel fetches to these APIs at runtime to grab the top ~15 abstracts.
Local Data: Clinical guidelines are stored locally in SQLite and retrieved via full-text search (FTS) ensuring exact matches on medical terminology.
Inference: I’m using Gemini 2.5 Flash to process the concatenated abstracts. The massive context window allows me to feed it distinct search results and force strict citation mapping without latency bottlenecks.
Workflow Tools (The "Integration"): I also built a "reasoning layer" to handle complex patient histories (Case Mode) and draft documentation (SOAP Notes). Case Mode Demo (GIF): https://imgur.com/a/h01Zgkx https://imgur.com/a/h01Zgkx Note Gen Demo (GIF): https://imgur.com/a/DI1S2Y0 https://imgur.com/a/DI1S2Y0
Why no Vector DB? In medicine, "freshness" is critical. If a new trial drops today, a pre-indexed vector store might miss it. My real-time approach ensures the answer includes papers published today.
Business Model: The clinical search is free. I plan to monetize by selling billing automation tools to hospital admins later.
Feedback Request: I’d love feedback on the retrieval latency (fetching live APIs is slower than vector lookups) and the accuracy of the synthesized answers.
- deleted 9mo ago[deleted]
- neil_naveen 9mo agoFYI, You are using Clerk in development mode
- amber_raza 9mo agoOof, good catch! I must have left the test keys active in the deployment config. Swapping them to production keys right now. Thanks for the heads up!
- bflesch 9mo agoSomehow "clerk" is on my ublock origin blocklist and therefore the whole website is not loading. I didn't add "clerk" to the blocklist so it must've been added by one of the blocklists that ublock origin is subscribed to, so there must be a good reason why "clerk" is on that blocklist. When building a product for medical audience which might care a lot about privacy maybe don't use components which are shady enough that they end up on blocklists. Edit: > Why no Vector DB? In medicine, "freshness" is critical. If a new trial drops today, a pre-indexed vector store might miss it. My real-time approach ensures the answer includes papers published today. This is total rubbish - did you talk to a single medical practitioner when building this? Nobody will do new treatments on their patients if a new paper was "published" (whatever that means, just being added to some search index). These people require trusted source, experimental treatment is only done for private clients who have tried all other options.
- amber_raza 9mo agoThanks for the feedback—this is helpful. 1. Re: Clerk/uBlock: You were spot on. The default Clerk domain often gets flagged by strict blocklists. I just updated the DNS records to serve auth from a first-party subdomain (clerk.getevidex.com) to resolve this. It should be working now. 2. Re: Freshness & 'Rubbish': You are absolutely right that standard of care doesn't (and shouldn't) change overnight based on one new paper. However, the decision to ditch the Vector DB for Live Search wasn't about pushing 'experimental treatments'—it was about Safety and Engineering constraints: Retractions & Safety Alerts: A stale vector index is a safety risk. If a major paper is retracted or a drug gets a black-box warning today, a live API call to PubMed/EuropePMC reflects that immediately. A vector store is only as good as its last re-index. The 'Long Tail': Vectorizing the entire PubMed corpus (35M+ citations) is expensive and hard to keep in sync. By using the search APIs directly, we get the full breadth of the database (including older, obscure case reports for rare diseases) without maintaining a massive, potentially stale index. The goal isn't to be 'bleeding edge'—it's to be 'currently accurate'.
- breadislove 9mo agoa good system (like openevidence) indexes every paper released and semantic search can incredible helpful since the the search api of all those providers are extremely limited in terms of quality. now you get why those system are not cheap. keeping indexes fresh, maintaining high quality at large scale and being extremely precise is challenging. by having distributed indexes you are at the mercy of the api providers and i can tell you from previous experience that it won't be 'currently accurate'. for transparency: i am building a search api, so i am biased. but i also build medical retrieval systems for some time.
- adit_ya1 9mo agoOut of curiosity, what's the prioritization of evidence (RTC Metanalysis > RTC > observational ) etc, and what's the end user benefit over a tool like OpenEvidence? You mention that other tools are expensive, slow, or increasingly heavy with pharma ads, but OpenEvidence for now seems to be pretty similiar with offerings, speed, and responses. What's your pitch as to why one should prefer this?
- amber_raza 9mo agoGreat questions. 1. Prioritization: I instruct the model to prioritize evidence in this hierarchy: Meta-Analyses & Systematic Reviews > RCTs > Observational Studies > Case Reports. It explicitly deprioritizes non-human studies unless specified. 2. Why not OpenEvidence? OE is excellent! But we made two architectural choices to solve different problems: 'Long Tail' Coverage: OE relies on a pre-indexed vector store, which often creates a blind spot for niche/rare diseases where papers aren't in the 'Top 1% of Journals.' Because Evidex queries live APIs, we catch the obscure case reports that static indexes often prune out. Workflow: OE is a 'Consultant' (Q&A). Evidex is a 'Resident' (Grunt work). The 'Case Mode' is built to take messy patient histories and draft the actual documentation (SOAP Notes/Appeals) you have to write after finding the answer.
- eoravkin 9mo agoOut of curiosity, did you actually see any pharma ads on OpenEvidence?
- amber_raza 9mo agoGreat question. I haven't seen banner ads on OpenEvidence yet, but the 'hidden tax' of free tools is often Publisher Bias. Users have noted that some current tools heavily overweight citations from 'Partner Journals' (like NEJM/JAMA) because they index the full text, effectively burying better papers from non-partner journals in the vector retrieval. My goal is strictly Neutral Retrieval. By hitting the PubMed/OpenAlex APIs live, Evidex treats a niche pediatric journal with the same relevance weight as a major publisher, ensuring the 'Long Tail' of evidence isn't drowned out by business partnerships.
- breadislove 9mo agothis might be interesting: https://www.theinformation.com/articles/chatgpt-doctors-startup-doubles-valuation-12-billion-revenue-surges https://www.theinformation.com/articles/chatgpt-doctors-star... > $150M RR on just ads, +3x from August. On <1M users. source: https://x.com/ArfurRock/status/1999618200024076620 https://x.com/ArfurRock/status/1999618200024076620
- amber_raza 9mo agoWhoa. $150M ARR on ads is a wild stat. Thanks for sharing that source. It really validates the thesis that unless the user pays (SaaS), the Pharma companies are the real customers.
- eoravkin 9mo agoYou built a cool product. I'm actually one of the founders of https://medisearch.io https://medisearch.io which is similar to what you are building. I think the long-tail problem that you describe can be solved in other ways than with live APIs and you may find other problems with using live APIs.
- dataviz1000 9mo agoI'm working on building an AI agent that creates queries over a time-series database focused on financial data. For example, it can quantify Federal Reserve reports and generate a table showing how SPY reacted 30 minutes after, at EoD, at the next day’s open, and at the next day’s EoD. It will plan the database query and then query the data from a materialized view. It is magic! How would biomedical researchers use tons of time-series data? A better question is: what questions are biomedical researchers asking with time-series data? I'm a lot more interested in generalized querying over time-series data than just financial data. What would be a great proof of concept?
- amber_raza 9mo agoThat sounds like a fascinating project. To answer your question: In the biomedical world, the 'Time-Series' equivalent is Patient Telemetry (Continuous Glucose Monitors, ICU Vitals, Wearables). The Question Researchers Ask: 'Can we predict sepsis/stroke 4 hours before it happens based on the velocity of change in Heart Rate + BP?' Right now, Evidex is focused on the Unstructured Text (Literature/Guidelines) rather than the structured time-series data, but the 'Holy Grail' of medical AI is eventually combining them: Using the Literature to interpret the Live Vitals in real-time.
- vikas-sharma 9mo ago[dead]
- jph 9mo agoGreat project. Want to contact me when you'd like to talk? I do software engineering for clinicians at a health care organization, and I'd love to have my teams try your work in their own contexts. Email joel@joelparkerhenderson.com.
- amber_raza 9mo agoThanks, Joel! This is exactly the kind of clinical workflow I built 'Case Mode' for. I will send you an email shortly to get connected. I'd love to get your teams set up with a pilot instance. Appreciate the reach out.
- OutOfHere 9mo agoAll such custom sites are increasingly unnecessary since modern thinking AIs like ChatGPT 5.2 Extended and Gemini 3 Pro do an incredible job surfacing good papers. In my experience, the benefit comes from using multiple AIs because they all have blind spots, and none is pareto optimal. As a patient, sometimes I don't want the AI to have my entire medical history, as this lets me consider things from different angles. For each chat, I give it the reconstructed history that I think is sufficient. I want it to be an explorer more than a doctor.
- amber_raza 9mo agoThat is a fair critique. The frontier models are getting incredible at general reasoning. The gap Evidex fills isn't 'Intelligence'. It is Provenance and Liability. Strict Sourcing: Even advanced models can hallucinate a plausible-sounding study. Evidex constrains the model to answer only using the abstracts returned by the API. This reduces the risk of a 'creative' citation. Explorer vs. Operator: You mentioned using AI as an 'explorer' (Patient use case). Doctors are usually 'operators'. They need to find the specific dosage or guideline quickly to close a chart. I view this less as replacing Gemini/GPT. It is more of a 'Safety Wrapper' around them for a high-stakes environment.
- OutOfHere 9mo agoThe problem is that doctors almost always, except perhaps in the emergency department, are currently too full of themselves, and are not open to reading relevant research unless a patient like me forces it upon the doctor. Maybe they are busy but that doesn't work for the patient. Even upon such forcing of the patient sharing research, the doctor will often read only a single line from an entire paper. How do you change this culture? It doesn't serve the patient too well to get an inaccurate root cause diagnosis from the doctors as I often do. It comes upon the patient to really spend the time investigating and testing hypotheses and theories, failing which the root causes go ignored, and one ends up taking too many unnecessary or even harmful pharmaceuticals.
- amber_raza 9mo ago
- pdyc 9mo agoI like your approach of "smart routing" but using regex/keywords based approach has a problem that it does not captures semantic similarity of keywords so search with similar intents are missed, how are you handling it? or you dont need to handle it since it is for domain experts and they are likely to search based on keywords(dictionary)?
- amber_raza 9mo agoYou hit the nail on the head regarding the 'semantic gap'. Currently, I handle this via Smart Routing. The engine analyzes the intent of your query (e.g. identifying if you’re looking for an RCT, a specific guideline, or drug dosing) and routes it to the most relevant clinical database using high-precision keyword matching. I chose this deterministic approach for the launch to ensure clinical precision. While vector/semantic search is great for general concepts, it can sometimes surface 'similar-ish' papers that miss the specific medical nuances (like a specific ICD-10 code or dosage) required for clinical evidence. The LLM (Gemini 2.5 Flash) currently lives in the Synthesis Layer. It takes the raw, high-precision results and synthesizes them into the clinical summaries you see. I actually have LLM-based query expansion (translating natural language into robust MeSH/Boolean strings) built into the infrastructure, but I am keeping it in 'staging' right now. I want to ensure that as I bridge that semantic gap, I don't sacrifice the deterministic accuracy that medical professionals expect.
- craigdalton 9mo agoExcuse the blunt metaphor, but there is a risk here of turning on a fire-hose of "fresh" garbage. John Ioannidis, one of the doyens of evidence based medicine very persuasively argues - Why Most Published Research Findings Are False https://pmc.ncbi.nlm.nih.gov/articles/PMC1182327/ https://pmc.ncbi.nlm.nih.gov/articles/PMC1182327/ That is why platforms pay physicians/epidemiologists/ specialists in their field hundreds of dollars per hour to sort the good from bad papers. After my training as a doctor I did a Masters in Clinical Epidemiology and spent an afternoon each week in a tutorial that reviewed papers in the top journals - about 20-30% of them had major flaws that were either ignored or dismissed by the authors. It may be worse now. LLMs still have trouble picking up the subtleties of medical science and will miss papers with major flaws. I just did a test on a paper that is often quoted as providing evidence of excess cancer risk in communities living close to unconventional gas facilities. When I asked ChatGPT 5.2 to review the pape for evidence of increase cancer risk with a simple prompt it said the paper found such a risk. However, when I wrote a multi-discipline based prompt for 5.2 and Gemini 3 pro, it found the fatal flaw in the paper and advised it did not provide evidence. See the prompt and consider how the prompts would have to be individually developed for each paper and meta-analysis. For review of meta-analysis you would need prompts developed by expert methodologists and discipline specialists- here is the prompt that worked: You are an environmental epidemiologist and exposure scientist, critially review this papers claim that the measured levels of unconventional gas emissions provide evidence of excess cancer risk: https://link.springer.com/article/10.1186/1476-069X-13-82 https://link.springer.com/article/10.1186/1476-069X-13-82
- amber_raza 9mo agoThis is a fantastic critique. Spot on. Freshness without appraisal is just an accelerated firehose of noise. 1. The Garbage Filter: Right now, I rely on a strict Hierarchy of Evidence to mitigate this (prioritizing Cochrane/Meta-analyses over observational studies), but you are absolutely right that LLMs can miss fatal methodological flaws in a single, high-ranking paper. 2. The 'Critic' Agent: I’m currently experimenting with a secondary 'Critic' pass. This is an LLM agent specifically prompted to act as a skeptic/methodologist to flag limitations before the main synthesis happens. 3. Multi-discipline prompting: The prompt you provided is a great case study in persona-based auditing. I’d love to learn more about the specific 'disciplines' or archetypes you’ve found most effective at catching these flaws. That is exactly the kind of domain expertise I’m trying to encode into the system.