5 ms·
I just tried this in Claude Code. I made an MCP server whose tool output is declared as an integer but it returns a string at runtime. Claude Code validated th
by cle 1y ago
I just tried this in Claude Code. I made an MCP server whose tool output is declared as an integer but it returns a string at runtime.
Claude Code validated the response against the schema and did not pass the response to the LLM.
test - test_tool (MCP)(input: "foo")
⎿ Error: Output validation error: 'bar' is not of type 'integer'
- ohdeargodno 1y agoThis time. Can you guarantee it will validate it every time ? Can you guarantee the way MCPs/tool calling are implemented (which is already an incredible joke that only python brained developers would inflict upon the world) will always go through the validation layer, are you even sure of what part of Claude handles this validation ? Sure, it didn't cast an int into a Toyota Yaris. Will it cast "70Y074" into one ? Maybe a 2022 one. What if there are embedded parsing rules into a string, will it respect it every time ? What if you use it outside of Claude Code, but just ask nicely through the API, can you guarantee this validation still works ? Or that they won't break it next week ? The whole point of it is, whichever LLM you're using is already too dumb to not trip when lacing its own shoes. Why you'd trust it to reliably and properly parse input badly described by a terrible format is beyond me.
- cle 1y agoThis is deterministic, it is validating the response using a JSON Schema validator and refusing to pass it to an LLM inference. I can't gaurantee that behavior will remain the same more than any other software. But all this happens before the LLM is even involved. > The whole point of it is, whichever LLM you're using is already too dumb to not trip when lacing its own shoes. Why you'd trust it to reliably and properly parse input badly described by a terrible format is beyond me. You are describing why MCP supports JSON Schema. It requires parsing & validating the input using deterministic software, not LLMs.
- whoknowsidont 1y ago>This is deterministic, it is validating the response using a JSON Schema validator and refusing to pass it to an LLM inference. No. It is not. You are still misunderstanding how this works. It is "choosing" to pass this to a validator or some other tool, _for now_. As a matter of pure statistics, it will simply not do this at some point in the future on some run. It is inevitable.
- cle 1y agoI'd encourage you to read the MCP specification: https://modelcontextprotocol.io/specification/2025-06-18/server/tools#output-schema https://modelcontextprotocol.io/specification/2025-06-18/ser... Or write a simple MCP server and a client that uses it. FastMCP is easy: https://gofastmcp.com/getting-started/quickstart https://gofastmcp.com/getting-started/quickstart You are quite wrong. The LLM "chooses" to use a tool, but the input (provided by the LLM) is validated with JSON Schema by the server, and the output is validated by the client (Claude Code). The output is not provided back to the LLM if it does not comply with the JSON Schema, instead an error is surfaced.
- whoknowsidont 1y agoWhy do you think anything you said contradicts what I'm saying? I promise you I'm probably far more experienced in this field than you are. >The LLM "chooses" to use a tool Take a minute to just repeat this a few times.
- fauigerzigerk 1y agoMCP requires that servers providing tools must deterministically validate tool inputs and outputs against the schema. LLMs cannot decide to skip this validation. They can only decide not to call the tool. So is your criticism that MCP doesn't specify if and when tools are called? If so then you are essentially asking for a massive expansion of MCP's scope to turn it into an orchestration or workflow platform.
- dragonwriter 1y ago
- dragonwriter 1y ago> Can you guarantee it will validate it every time ? Yes, to the extent you can guarantee the behavior of third party software, you can (which you can't really guarantee no matter what spec the software supposedly implements, so the gaps aren't an MCP issue), because “the app enforces schema compliance before handing the results to the LLM” is deterministic behavior in the traditional app that provides the toolchain that provides the interface between tools (and the user) and the LLM, not non-deterministic behavior driven by the LLM. Hence, “before handing the results to the LLM”. > The whole point of it is, whichever LLM you're using is already too dumb to not trip when lacing its own shoes. Why you'd trust it to reliably and properly parse input badly described by a terrible format is beyond me. The toolchain is parsing, validating, and mapping the data into the format preferred by the chosen models promot template, the LLM has nothing to do with doing that, because that by definition has to happen before it can see the data. You aren't trusting the LLM.
- whoknowsidont 1y ago>The toolchain is parsing, validating, and mapping the data into the format preferred by the chosen models promot template, the LLM has nothing to do with doing that The LLM has everything to do with that. The LLM is literally choosing to do that. I don't know why this point keeps getting missed or side-stepped. It WILL, at some point in the future and given enough executions, as a matter of statistical certainty, simply not do that above, or pretend to do the above, or do something totally different at some point in the future.
- dragonwriter 1y ago> The LLM has everything to do with that. The LLM is literally choosing to do that. No, the LLM doesn't control on a case-by-caae basis what the toolchain does between the LLM putting a tool call request in an output message and the toolchain calling the LLM afterwards. If the toolchain is programmed to always validate tool responses against the JSON schema provided by MCP server before mapping into the LLM prompt template and calling the LLM again to handle the response, that is going to happen 100% of the time. The LLM doesn't choose it. It CAN'T because the only way it even knows that the data has come back from the tool call is that the toolchain has already done whatever it is programmed to do, ending with mapping the response into a prompt and calling the LLM again. Even before MCPs or even models specifically trained and with vendor-provided templates for tool calling (but after the ReAct architecture was described), it was like a weekend project to implement a basic framework supporting tooling calling around a local or remote LLM. I don't think you need to do that to understand how silly the claim that the LLM controls what the toolchain does with each response and might make it not validate it is, but certainly doing it will give you a visceral understanding of how silly it is.
- thwarted 1y agoAs an example. "1979010112345" is a unix timestamp that looks like it might be Jan 1 1979 datetime formatted as an integer, but is really Sep 17 2032 05:01:52.
- whoknowsidont 1y agoHow many times does this need to be repeated. It works in this instance. On this run. It is not guaranteed to work next time. There is a error percentage here that makes it _INEVITABLE_ that eventually, with enough executions, the validation will pass when it should fail. It will choose not to pass this to the validator, at some point in the future. It will create its own validator, at some point in the future. It will simply pretend like it did any of the above, at some point in the future. This might be fine for your B2B use case. It is not fine for underlying infrastructure for a financial firm or communications.
- cle 1y agoEvery time the LLM uses this tool, the response schema is validated--deterministically. The LLM will never see a non-integer value as output from the tool.
- whoknowsidont 1y agoCan you please diagram out, using little text arrows ("->"), what you think is happening so I can just fill in the gap for you?
- masafej536 1y agollm tool call -> mcp client validates the schema -> mcp client calls the tool -> mcp server validates the schema -> mcp server responds with the result -> mcp client passes the tool result into llm
- redandblack 1y agonot a developer. what happens if this schema validation fails here - what will the mcp server respond with and what will the llm do next (in a deterministic sense)? llm tool call -> mcp client validates the schema -> mcp client calls the tool -> mcp server validates the schema
- 1y ago