MEASURED
Finding · verify
At the sample size used, the question the experiment was built to answer could not be answered either way, and that is reported rather than rounded to a result.
§ 1
The record
FINDING · judge-agreement-undetermined-at-n45verify
At the sample size used, the question the experiment was built to answer could not be answered either way, and that is reported rather than rounded to a result.
- Claim
- At 45 answers, a judge with zero disagreements has an agreement interval spanning its whole range; the perfect score reads as 'no counterexample found at this sample size', not as demonstrated reliability.
- Status
- MEASURED Empirical, with a stated sample.
- Question
- Passed is not correct
- Subject
- An ensemble judge on 45 GSM8K answers (existing models). third-party subject
- Frame
- 45 answers; a zero-disagreement judge.
- Method
- Agreement interval computed at the observed sample size.
- Oracle
- The GSM8K answer key.
- Negative control
- Not applicable The result is that the interval is degenerate; there is nothing for a control to distinguish.
- Denominator
- 0 disagreements in 45; interval spans the whole range.
- Limitation
- That the ensemble judge is reliable. A degenerate interval at the endpoints cannot express uncertainty, and at larger scale even ensemble judges are imperfect.
- Source
- repowazdogz-droid/evaltrust @ 18e999a5 archivedRepository archived 2 September 2026 after a dated correction to its second experiment. The numbers cited here are from the first experiment and are not affected by the correction, which withdraws a vendor-level interpretation of a separate contrast. Read the correction.
- Reproduce
python scripts/build_report.py && python scripts/recompute.py && pytest -q
- Independent reproduction
- None known.
§ 2
Where this sits
This finding answers Passed is not correct and supports the VERIFY stage of the operating method.