27 ms·
Show HN: I built an 11-LLM consensus engine to detect AI hallucination
- jaquelinejaque 3mo ago[flagged]
- Lionga 3mo ago[flagged]
- jaquelinejaque 3mo ago[flagged]
- dang 3mo agoPlease don't post snarky or aggressive comments, and especially not in Show HN threads. You broke both the site guidelines and the Show HN guidelines badly here. If you'd please review https://news.ycombinator.com/showhn.html https://news.ycombinator.com/showhn.html and https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and stick to the rules when commenting, we'd appreciate it. We've had to ask you this at least once before: https://news.ycombinator.com/item?id=44759301 https://news.ycombinator.com/item?id=44759301. Edit: actually, the commenting history of this account is breaking the site guidelines so badly and so often that I've banned it. Please see https://news.ycombinator.com/item?id=48611655 https://news.ycombinator.com/item?id=48611655.
- jmtrevarton 3mo agoDoes the user set up API keys for those 11 LLMs or is API cost included in the product? Do you test for tool hallucination or only information hallucination?
- jaquelinejaque 3mo ago[flagged]
- r0fl 3mo agoInteresting idea I get codex to use openrouter api and ask it to find 5 cheap but highly efficient LLMs at the task that km doing based on benchmarks and descriptions I then run the query through all 5, get a markdown file for each in case I want to read through it later and have codex analyze and improve things based on those 5 outputs It’s very easy and can scale to 11 or more LLMs with the same api
- jaquelinejaque 3mo ago[flagged]
- jaquelinejaque 3mo ago[flagged]
- pedromlsreis 3mo agoGenuine question, how did you get to these 11 LLMs instead of 10 or 12? I'm interested in understanding how you did benchmark these 11 LLMs or whether it was an arbitrary ensemble you selected.
- madikz 3mo ago[flagged]