3 ms·
I built a protocol to test whether small local LLMs (1B-31B, running on Ollama) can recognize and reason about their own prior outputs. Give a model a blank pag
by daniel-navarro 5mo ago
I built a protocol to test whether small local LLMs (1B-31B, running on Ollama) can recognize and reason about their own prior outputs. Give a model a blank page, let it write three entries, then ask it structured questions about what it wrote. Score the answers with three independent AI judges from different labs (GPT, Gemini, Sonnet).
Two things I didn't expect:
1. The three judges don't agree. GPT-5.4-mini almost never uses "1" on a 0-4 rubric -- it collapses the scale, redirecting that mass to "2". Gemini and Sonnet use the full scale normally. Same rubric, same responses, different scores. A forensic audit of all 87 scored runs traces the mechanism: GPT resolves an ambiguity in the scoring prompt differently from the other two. If you're using a single AI judge in your evals, you're measuring the judge as much as the subject.
2. Across three protocol iterations, I discovered the instrument was driving most of the measured behavior. Bootstrap prompts primed model outputs. Default thinking-mode settings silenced some models entirely (they spent all their tokens thinking and produced empty responses). Fixing these confounds reversed earlier conclusions. The models were fine -- the microscope was dirty.
Everything is open: runnable instrument, all raw data (2,349 judge records), the working paper, and the forensic analysis. MIT license.