Findings · Evaluation integrity
At 45 answers, a judge with zero disagreements has an agreement interval spanning its whole range; the perfect score reads as 'no counterexample found at this sample size', not as demonstrated reliability.
Evidence
An ensemble judge that never disagrees with the gold answer over forty-five items looks perfect. The agreement interval is degenerate at the endpoints, by construction: with no disagreements there is nothing to be uncertain about, so the interval collapses and stops carrying information. The honest reading is “no counterexample found at this sample size,” not “proven reliable.” The report says so where the number appears, and cites a larger-scale study where even ensemble judges fall short.
Sample. 45 real GSM8K trajectories. The ensemble and numeric-oracle judges have zero disagreements, so the interval is degenerate; the substring judge's interval spans its full range on 2 of 45 disagreements.
Boundary
That the ensemble judge is reliable. A degenerate interval at the endpoints cannot express uncertainty, and at larger scale even ensemble judges are imperfect.
Independent reproduction. None known.
Reproduction
python scripts/build_report.py && python scripts/recompute.py && pytest -q - Toolchain
- Python >=3.11, numpy >=1.24, scipy >=1.10
- Expected output
- writeup/report.md rebuilt; digest byte-identical to 62bb20c4a8bcd334a6ccd2d04398e6bff29943f76916ad4c5a92c51a723dc10a
- Claim stated at
- repowazdogz-droid/evaltrust · writeup/report.md:38,50,110
- Verified at
- repowazdogz-droid/evaltrust@18e999a5 (2026-07-20)
This pass. Confirmed by source read at 18e999a5 (committed digest present); the recompute pipeline was not re-run this pass.
Artifact: repowazdogz-droid/evaltrust