2 ms·
Cool idea — “book as rubric” is a nice way to avoid vague self-critiques. One thing I’ve found helps keep agent review loops from getting shallow is to separat
by YaraDori 7mo ago
Cool idea — “book as rubric” is a nice way to avoid vague self-critiques.
One thing I’ve found helps keep agent review loops from getting shallow is to separate levels of critique:
1) a fast “lint” pass (format, obvious bugs, missing tests)
2) a domain pass that’s forced to cite specific passages/rules from the rubric
3) a “counterexample” pass: reviewer must propose at least 1 concrete failing scenario + how to reproduce
Question: are you capturing the reviewer’s evidence (links, excerpts, failing cases) in a structured log so a human can audit why the agent changed something?
Related: I’m working on SkillForge (https://skillforge.expert https://skillforge.expert) — we’ve been thinking about similar auditability, but for UI workflows: record once, then replay with checkpoints + retries, so humans approve at meaningful boundaries instead of every click.