3 ms·
love the direction, but the problem for me has been about creating stronger evals and knowing what I should be checking for. does this help me understand that?
by vigjam 16d ago
love the direction, but the problem for me has been about creating stronger evals and knowing what I should be checking for. does this help me understand that?
- edstack 16d ago[dead]
- prathmeshmcp 16d agoyeah, in my blog on Effective MCP (https://www.mcpjam.com/blog/effective-mcp-part-1 https://www.mcpjam.com/blog/effective-mcp-part-1) I introduce this framework called the "User-Value Chain" that essentially breaks down the full client-server request flow into distinct stages. You want to test the full request flow between client and server: connection, tool discovery, tool selection, calls, responses, and whether the user’s GOAL was actually achieved. What's hard: an external agent sits between your user and your MCP server, in a client you don't control (ChatGPT, claude etc.). Your tests need to cover how that client and agent find and uses your tools. Recommend checking for : - (deterministic assertions + non-deterministic judge checks) at essentially every stage of the request flow from client->server (User Value Chain) - across many clients where YOUR target users are at (unfort. these clients change behavior every other day) and across harness + models - our default eval assertions cover things like input schemas, valid arguments, repeated calls, errors, latency, and response size; pair those deterministic checks with non-deterministic evals for tool choice, response interpretation, and task completion.
- synfrac 14d ago[flagged]