7 ms·
Show HN: Ask-human-mcp – zero-config human-in-loop hatch to stop hallucinations
While building my startup i kept running into the issue where ai agents in cursor create endpoints or code that shouldn't exist, hallucinates strings, or just don't understand the code.
ask-human-mcp pauses your agent whenever it’s stuck, logs a question into ask_human.md in your root directory with answer: PENDING, and then resumes as soon as you fill in the correct answer.
the pain:
your agent screams out an endpoint that never existed
it makes confident assumptions and you spend hours debugging false leads
the fix:
ask-human-mcp gives your agent an escape hatch. when it’s unsure, it calls ask_human(), writes a question into ask_human.md, and waits. you swap answer: PENDING for the real answer and it keeps going.
some features:
- zero config: pip install ask-human-mcp + one line in .cursor/mcp.json → boom, you’re live
- cross-platform: works on macOS, Linux, and Windows—no extra servers or webhooks.
- markdown Q\&A: agent calls await ask_human(), question lands in ask_human.md with answer: PENDING. you write the answer, agent picks back up
- file locking & rotation: prevents corrupt files, limits pending questions, auto-rotates when ask_human.md hits ~50 MB
the quickstart
pip install ask-human-mcp
ask-human-mcp --help
add to .cursor/mcp.json and restart:
{
"mcpServers": {
"ask-human": { "command": "ask-human-mcp" }
}
}
now any call like:
answer = await ask_human(
"which auth endpoint do we use?",
"building login form in auth.js"
)
creates:
### Q8c4f1e2a
ts: 2025-01-15 14:30
q: which auth endpoint do we use?
ctx: building login form in auth.js
answer: PENDING
just replace answer: PENDING with the real endpoint (e.g., `POST /api/v2/auth/login`) and your agent continues.
link:
github -> https://github.com/Masony817/ask-human-mcp https://github.com/Masony817/ask-human-mcp
feedback:
I'm Mason a 19yo solo-founder at Kallro. Happy to hear any bugs, feature requests, or weird edge cases you uncover - drop a comment or open an issue!
buy me a coffee -> coff.ee/masonyarbrough
- throwaway314155 1y agoNot certain that your definition of hallucination matches mine precisely. Having said that, this is so simple yet kinda brilliant. Surprised it's not a more popular concept already.
- loloquwowndueo 1y ago- someone sets up an “ask human as a service mcp” - demand quickly outstrips offer of humans willing to help bots - someone else hooks up AI to the “ask human saas” - we now have a full loop of machines asking machines
- TZubiri 1y agoThis is pretty much already possible in any economy, but quite a waste. Not much is stopping you from buying products from a retailer and selling them at a wholesaler, but you'd lose money in doing so.
- 4ndrewl 1y agoI mean, losing money is practically de rigeur in the startup community right?
- olalonde 1y agoI built this - but mostly as a joke / proof-of-concept: https://github.com/olalonde/mcp-human https://github.com/olalonde/mcp-human
- aziaziazi 1y agoCool project! Naive question: does mechanical turk uses llm now?
- lordmauve 1y agoFinally, the "AI" turns out to be 700 Indians. We now have the full loop of humans asking machines asking humans pretending to be machines. Civilisation collapses
- franky47 1y agoAI stands for Actual Indians.
- 1y ago
- conception 1y agoWhat sort of prompt are you using for this?
- kordlessagain 1y agoThe prompt is (mostly) built using the tool loads in the MCP server. In Python, the @mcp.tool() decorators provide the context of tool to the prompt, which is then submitted (I believe) with each call to the LLM.
- rgbrenner 1y agoSounds similar to `ask_followup_question` in Roo
- kjhughes 1y agoCool conceptually, but how exactly does the agent know when it's unsure or stuck?
- Groxx 1y agoThe same way it knows anything else. So not at all, but that doesn't mean it's not useful.
- TZubiri 1y agoSo we are just pushing the issue to another, less debuggable layer. Cool.
- kjhughes 1y agoI'll try to give you credit for more than dismissing my question off-hand... Yes, it may not need to know with perfect certainty when it's unsure or stuck, but even to meet a lower bar of usefulness, it'll need at least an approximate means of determining that its knowledge is inadequate. To purport to help with the hallucination problem requires no less. To make the issue a bit more clear, here are some candidate components to a stuck() predicate: - possibilities considered - time taken - tokens consumed/generated (vs expected? vs static limit? vs dynamic limit?) If the unsure/stuck determination is defined via more qualitative prompting, what's the prompt? How well has it worked?
- Groxx 1y agoI don't believe[1] any of those are part of the MCP protocol - it's essentially "the LLM decided to call it, with X arguments, and will interpret the results however it likes". It's an escape hatch for the LLM to use to do stuff like read a file, not a monitoring system that acts independently and has control over the LLM itself. (But you could build one that does this, and ask the LLM to call it and give your MCP that data... when it feels like it) So you'd be using this by telling the LLM to run it when it thinks it's stuck. Or needs human input. 1: I am not anything even approaching deeply knowledgeable about MCP, so please, someone correct me if I'm wrong! There do seem to be some bi-directional messaging abilities, e.g. notification, but to figure out thinking time / token use / etc you would need to have access to the infrastructure running the LLM, e.g. Cursor itself or something.
- mgraczyk 1y agoIf you are answering these questions yourself, why not just add something like this to your cursor rules? "If you don't know the answer to a question and need the answer to continue, ask me before continuing" Will you have some other person answer the question?
- deadbabe 1y agoHaving another person answer the question is pretty much the obvious route this will go.
- mgraczyk 1y agoBut then that means they are editing a markdown file on your computer? How is that meant to work? I like the idea but would rather it use Slack or something if it's meant to ask anyone.
- echollama 1y agothis is mainly meant as a way to conversate with the model while you are programming with it. This is not meant to pull questions to a team but more to pair program. a markdown file is best for syntax in an llm prompt and also just easiest to have open and answer questions with. If i had more time and could i would build an extension into cursor.
- mgraczyk 1y agoWhy not have the model ask in the chat? It's a lot easier to just talk to it than open a file. The article mentions cursor so it sounds like you're already using cursor?
- echollama 1y agowould probably work better, this is just how i threw it together as an internal tool a long time ago. i just improved it and shipped it to opensource it.
- superb_dev 1y agoThis site is impossible to read on my phone. Part of the left side of the screen is cut off and I can’t scroll it into view
- threeseed 1y ago> an mcp server that lets the agent raise its hand instead of hallucinating a) It doesn't know when it's hallucinating. b) It can't provide you with any accurate confidence score for any answer. c) Your library is still useful but any claim that you can make solutions more robust is a lie. Probably good enough to get into YC / raise VC though.
- echollama 1y agoreasoning models know when they are close to hallucinating because they are lacking context or understanding and know that they could solve this with a question. this is a streamlined implementation of a interanlly scrapped together tool that i decided to open-source for people to either us or build off of.
- geraneum 1y ago> reasoning models know when they are close to hallucinating because they are lacking context or understanding and know that they could solve this with a question. I’m interested. Where can I read more about this?
- threeseed 1y ago> reasoning models know when they are close to hallucinating because they are lacking context or understanding and know that they could solve this with a question You've just described AGI. If this were possible you could create an MCP server that has a continually updated list of FAQ of everything that the model doesn't know. Over time it would learn everything.
- xeonmc 1y agoUnless there is as yet insufficient data for meaningful answer.
- marshall300791 1y agohttps://arxiv.org/html/2407.14507v3 https://arxiv.org/html/2407.14507v3
- exclipy 1y agoWould be great if it pinged me on slack or whatsapp. I wouldn't notice if it simply paused waiting for the MCP call to return
- spacecadet 1y agoEasy enough to do with smolagents and fastmcp, its 20 lines of code.
- atoav 1y agoI am running an electronics/medialab in an university, the amount of fires bad electronics advice from LLMs caused already is probably non-zero. It is amazing how bad LLMs are when it comes to reasoning about simple dynamics within trivial electronic circuits and how eager they are to insist the opposite of how things work in the real world is the secured truth.
- spacecadet 1y agoIf the model responds with an obvious incorrect answer or hallucination, start over. Rephrase your input. Consider what output you are actually after... Adding to original shit output wont help you.
- ddalex 1y agoWhy wouldn't a rag-enabled ai be faster and better then humans at answering these documentation-grounded questions ?
- kordlessagain 1y agoThe same technique can be had by creating a "universal MCP tool" for the LLM to use if it thinks the existing tools aren't up to the job. The MCP language calls these "proxies".
- dkkergoog 1y ago[dead]
- PSBigBig 1y agoThanks for sharing this. Bookmarked!