MEASURED
Finding · verify
Two AI graders agreeing tells you less than it appears to, because part of that agreement is forced by both being right.
§ 1
The record
FINDING · judge-agreement-is-partly-forced-by-accuracyverify · convergence
Two AI graders agreeing tells you less than it appears to, because part of that agreement is forced by both being right.
- Claim
- Across 600 items and 171 judge pairs, judge errors correlated positively in 171 of 171 pairs at mean phi 0.567, and the correlation between pairwise agreement and accuracy fell from 0.824 to 0.077 once the algebraically forced component of agreement was removed.
- Status
- MEASURED Empirical, with a stated sample.
- Question
- Passed is not correct
- Subject
- Language-model judges on 600 items with a machine-checkable key (existing models). third-party subject
- Frame
- 600 items, 171 judge pairs; ground truth is a key, not human labelling.
- Method
- Pre-registered two-experiment study; the algebraically forced component of agreement is removed and the residual correlation reported.
- Oracle
- The machine-checkable answer key; report fields named per headline.
- Negative control
- Present An adversarial control (admissibility fixed before the data) fired as designed.
- Denominator
- 171 of 171 pairs positive, mean phi 0.567; corr 0.824 to 0.077.
- Limitation
- Not a claim that agreement between judges carries no information, and not a general law about language-model judges. It is a property of these judge checkpoints on these 600 items against a machine-checkable answer key. Ground truth here is a key, not human labelling, and the mechanism behind the second experiment is reported as open rather than settled.
- Source
- repowazdogz-droid/evaltrust @ c8196e1a archivedRepository archived 2 September 2026 after a dated correction to its second experiment. The numbers cited here are from the first experiment and are not affected by the correction, which withdraws a vendor-level interpretation of a separate contrast. Read the correction.
- Reproduce
python scripts/agreement/build_report.py && python scripts/agreement/sweep_writeup.py
- Independent reproduction
- None known.
§ 2
Where this sits
This finding answers Passed is not correct and supports the VERIFY stage of the operating method. It is an instance of the convergence mechanism.