Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
hidai25
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
hidai25
5mo ago
You’re right that I mixed runtime enforcement with CI drift/regression testing. Different layer, different job. I meant it as complementary, not equivalent. CrabTrap for runtime control, EvalView for deterministic testing/diffing.
2.
▲
by
hidai25
5mo ago
Interesting approach! I’ve been building something complementary on the deterministic side. LLM-as-judge guardrails are fundamentally probabilistic and can be gamed or hallucinate themselves (as several comments pointed out). That’s why I b
3.
▲
by
hidai25
9mo ago
Your agent worked yesterday. Today it's broken. What changed? EvalView catches regressions before they hit prod. Tool changes, hallucinations, cost spikes. Save a baseline. Run evalview run --diff. CI fails if behavior drifts.
4.
▲
Show HN: EvalView – Catch agent regressions before you ship (pytest for agents)
(github.com)
2 points
by
hidai25
9mo ago
|
1 comments
5.
▲
by
hidai25
10mo ago
Hi HN, I built EvalView after an agent that worked fine in dev started inventing numbers in prod. Tracing showed me what happened after the fact, but I wanted CI to fail the deploy the moment the agent drifted. EvalView is basically pytest
6.
▲
Show HN: EvalView pytest style tests for AI agents (budgets, hallucinations)
(github.com)
1 points
by
hidai25
10mo ago
|
1 comments