MEASURED
Finding · verify
Two off-the-shelf AI graders scoring the same work were wrong in opposite directions.
§ 1
The record
FINDING · judges-disagree-in-opposite-directionsverify
Two off-the-shelf AI graders scoring the same work were wrong in opposite directions.
- Claim
- Two language-model judges scoring the same 45 GSM8K answers disagreed with the ground truth in opposite directions: one systematically too strict, the other too lenient.
- Status
- MEASURED Empirical, with a stated sample.
- Question
- Passed is not correct
- Subject
- Two off-the-shelf language-model judges scoring 45 real GSM8K trajectories (existing models). third-party subject
- Frame
- 45 answers; per-answer labels are assistant-adjudicated; two of three judges also generated answers.
- Method
- Five graders score the same trajectories; every number recomputes from committed data under a digest check.
- Oracle
- The GSM8K integer answer key.
- Negative control
- None No control is claimed for this measurement; the sample is small and the labels are an authored scheme.
- Denominator
- gemma3 disagreed on 4 of 45, all too strict; llama3 on 2 of 45, all too lenient.
- Limitation
- That either judge is reliable, or that the direction generalises beyond this corpus. The sample is small; the per-answer correctness labels are assistant-adjudicated, not human; and two of the three judges also generated answers in the set, so self-preference is uncorrected.
- Source
- repowazdogz-droid/evaltrust @ 18e999a5 archivedRepository archived 2 September 2026 after a dated correction to its second experiment. The numbers cited here are from the first experiment and are not affected by the correction, which withdraws a vendor-level interpretation of a separate contrast. Read the correction.
- Reproduce
python scripts/build_report.py && python scripts/recompute.py && pytest -q
- Independent reproduction
- None known.
§ 2
Where this sits
This finding answers Passed is not correct and supports the VERIFY stage of the operating method.