Artifact
evaltrust
Two off-the-shelf LLM judges fail in opposite directions; a zero-disagreement judge's agreement interval is degenerate.
Archived
Archived 2 September 2026 after a dated correction: the second experiment's vendor-level interpretation was withdrawn; the first experiment's numbers stand and recompute from committed data. Read the correction.
What it establishes
That two off-the-shelf language-model judges disagree with the ground truth in opposite directions on a small GSM8K corpus, and that at that sample size a zero-disagreement judge's agreement interval is degenerate, measured, with digests that recompute offline.
What it does not establish
Judge reliability. The per-answer correctness labels are assistant-adjudicated, not human; two of the three judges also generated answers, so self-preference is uncorrected; the sample is small.
Method
Five graders, three language-model judges, a substring matcher, a numeric oracle, and an ensemble, score 45 real GSM8K trajectories, and every number recomputes from committed data under a digest check.
Results
gemma3 disagrees on 4 of 45, all too strict; llama3 on 2 of 45, all too lenient; the ensemble's agreement interval is degenerate at zero disagreements; the report digest is 62bb20c4.
What has to be trusted
The objective GSM8K integer answer key; the assistant-adjudicated per-trajectory labels, which are an authored coding scheme and part of the trusted base; the local Ollama judge models at temperature 0.
Prior work
Language-model-as-judge evaluation and inter-rater agreement. The contribution is the opposite-direction failure and the degenerate-interval honesty, with every number recomputing from committed data.
Reproduce it
python scripts/build_report.py && python scripts/recompute.py && pytest -q
Findings drawn from this artifact are on the evidence ledger.