3 ms·
Interesting approach! I’ve been building something complementary on the deterministic side. LLM-as-judge guardrails are fundamentally probabilistic and can be g
by hidai25 5mo ago
Interesting approach! I’ve been building something complementary on the deterministic side.
LLM-as-judge guardrails are fundamentally probabilistic and can be gamed or hallucinate themselves (as several comments pointed out).
That’s why I built EvalView — it does full trajectory snapshots + diffs so you can see exactly what changed, plus a lightweight zero-judge model-check that directly pings the model and reports drift level (NONE / WEAK / MEDIUM / STRONG).
Gives you deterministic regression detection that works alongside (or instead of) LLM judges.
https://github.com/hidai25/eval-view https://github.com/hidai25/eval-view
Curious how you handle drift detection in CrabTrap.
- pitched 5mo agoSecuring agents in real time and testing them for drift in CI are pretty different use-cases… This post is an AI-generated ad, isn’t it? It’s getting too hard to tell!
- hidai25 5mo agoYou’re right that I mixed runtime enforcement with CI drift/regression testing. Different layer, different job. I meant it as complementary, not equivalent. CrabTrap for runtime control, EvalView for deterministic testing/diffing. My bad on making it sound like a drive-by promo.
- Jerem-6ix 5mo ago[dead]