4 ms·
We build MCP servers that wrap fund APIs. The biggest performance variable we’ve found isn’t the model, it’s how much domain context the harness provides before
by prostheticrazor 8mo ago
We build MCP servers that wrap fund APIs. The biggest performance variable we’ve found isn’t the model, it’s how much domain context the harness provides before the model has to reason. Same model, generic prompt versus one loaded with our procedural docs - wider gap than switching between model generations. Which surprised me.
The post’s framing is right but undersells what the harness actually does in production. It’s your trust layer: what can the model touch, what can’t it, how cheaply do you recover when it gets something wrong. We spend something like 70% of engineering time on the recovery path, not the inference. Whether that ratio is right I’m not sure, but it’s where we’ve ended up.
On MCP overhead downthread: real, yes. In regulated environments you need the audit trail and the kill switch, and a tool boundary is how you get those. The unsolved part is keeping the protocol thin enough that you’re not burning tokens on ceremony.