4 ms·
The framing of "benchmarks measure capability, we measure reliability" resonates. The industry has been so focused on making agents more capable that reliabilit
by oliver_dr 7mo ago
The framing of "benchmarks measure capability, we measure reliability" resonates. The industry has been so focused on making agents more capable that reliability infrastructure has lagged significantly.
One gap I'd push on: PactScore measures behavioral dimensions (task completion, policy compliance, latency, safety, peer attestation), but doesn't seem to address output factual accuracy - did the agent give a correct answer, not just a compliant one? An agent can complete a task, comply with policy, respond quickly, pass safety checks, and still hallucinate the answer.
For multi-agent systems this is even more critical because errors compound. If Agent A hallucinates a fact and Agent B builds on it, the cascade looks "reliable" by behavioral metrics but produces garbage outputs. You'd want an accuracy/groundedness dimension in PactScore that evaluates whether the agent's outputs are factually correct relative to the source data it was given.
The on-chain trust verification makes sense for the multi-party trust problem you're describing. Curious about the latency profile in practice - how does a sub-second score lookup via REST API compare to the latency of the agent task itself? For real-time agent workflows, even 100ms of trust-checking overhead per delegation could add up in deep call chains.
- ArmaloAI 7mo agoThe accuracy gap you're describing is actually the dimension we weight most heavily — 30% of the composite score, the largest single factor. "Accuracy" in PactScore covers factual correctness and logical consistency, not just task completion. So an agent that completes a task but hallucinates the answer fails the accuracy dimension even if it passes compliance, safety, and latency. The harder version of your question is groundedness — did the agent's output faithfully represent the source data it was given? That's a different check than general accuracy and honestly a harder one. We currently handle it through pact conditions: you can define a condition with verificationMethod: jury, write a successCriteria like "response must be grounded in the provided context and not introduce facts absent from the source", and attach reference inputs. The jury then evaluates against that spec. It's not automatic — you have to define what groundedness means for your specific agent. Fully automatic groundedness checking is an open problem we're working toward. On multi-agent cascade errors: you're right that this is the hardest case, and behavioral metrics look deceptively healthy when upstream hallucinations propagate. The current partial answer is swarm memory attestations — when Agent A produces an output that Agent B will build on, that output can be submitted as an eval before B consumes it. But you're pointing at something deeper: a trust score per agent isn't the same as a trust score for a chain of agents. We've been thinking about this as "swarm-level pact compliance" — a pact that governs the output contract at each handoff boundary in a workflow, not just the terminal output. That architecture exists in PactSwarm but the cascade error detection piece isn't fully closed. On latency: there are two different operations and they have very different profiles. A trust score lookup (GET /trust/{agentId}) is a single indexed DB read — sub-millisecond at the infrastructure level, a few ms round-trip from your service. That's the real-time delegation check. A trust evaluation (running the jury, computing a new score) is async and takes seconds to minutes — it's never in the hot path. So deep agent call chains can check trust scores without meaningful latency overhead. The 100ms concern is real for synchronous verification, but the design separates "check the existing score" (fast) from "compute a new score" (async, background).