Findings · Evaluation integrity
Across 600 items and 171 judge pairs, judge errors correlated positively in 171 of 171 pairs at mean phi 0.567, and the correlation between pairwise agreement and accuracy fell from 0.824 to 0.077 once the algebraically forced component of agreement was removed.
Evidence
When two judges score the same answers, how often they agree is commonly read as a signal of how reliable they are. Part of that signal is arithmetic rather than evidence.
For a yes-or-no verdict against a yes-or-no answer key, two judges that are both wrong about an item have no choice but to agree about it. Agreement therefore has a floor that rises with accuracy, and a correlation between agreement and accuracy is partly created by that floor rather than observed in the judges. Removing the forced component leaves the part that is free to vary, which is the correlation between the judges’ errors. Measured here, that residual correlation is close to zero, while the errors themselves correlate positively in every pair examined.
This does not mean agreement is uninformative. It means a raw agreement statistic and an error-correlation statistic answer different questions, and only the second one is free to disagree with accuracy. A related stacked arrangement, in which one judge reviews another’s verdicts, caught fewer false passes rather than more.
An earlier pilot on 45 answers is kept on this site and is superseded by this study for anything about agreement and accuracy.
Sample. 600 items, 171 judge pairs, two task domains; intervals are judge-clustered rather than item-level, and the pre-registration and both amendments were committed before the data they govern existed.
Boundary
Not a claim that agreement between judges carries no information, and not a general law about language-model judges. It is a property of these judge checkpoints on these 600 items against a machine-checkable answer key. Ground truth here is a key, not human labelling, and the mechanism behind the second experiment is reported as open rather than settled.
Independent reproduction. None known. Every number re-derives from committed raw model outputs with no model access.
Reproduction
python scripts/agreement/build_report.py && python scripts/agreement/sweep_writeup.py - Toolchain
- Python; no model access required, the raw outputs are committed
- Expected output
- agreement_report.json carries mean_phi 0.567 with a judge-clustered interval, and the two correlation fields 0.824 and 0.077
- Claim stated at
- repowazdogz-droid/evaltrust · README.md, the claims-map table; writeup/agreement_report.json
- Verified at
- repowazdogz-droid/evaltrust@c8196e1a (2026-07-28)
This pass. Re-verified this pass by source read at commit c8196e1a; the report builder and the writeup sweep are in the repository and were not re-executed in this pass.
Artifact: repowazdogz-droid/evaltrust