Ω Omega Protocol Bring us a question
MEASURED

Finding · verify

At the sample size used, the question the experiment was built to answer could not be answered either way, and that is reported rather than rounded to a result.

§ 1

The record

FINDING · judge-agreement-undetermined-at-n45verify
At the sample size used, the question the experiment was built to answer could not be answered either way, and that is reported rather than rounded to a result.
Claim
At 45 answers, a judge with zero disagreements has an agreement interval spanning its whole range; the perfect score reads as 'no counterexample found at this sample size', not as demonstrated reliability.
Status
MEASURED Empirical, with a stated sample.
Question
Passed is not correct
Subject
An ensemble judge on 45 GSM8K answers (existing models). third-party subject
Frame
45 answers; a zero-disagreement judge.
Method
Agreement interval computed at the observed sample size.
Oracle
The GSM8K answer key.
Negative control
Not applicable The result is that the interval is degenerate; there is nothing for a control to distinguish.
Denominator
0 disagreements in 45; interval spans the whole range.
Limitation
That the ensemble judge is reliable. A degenerate interval at the endpoints cannot express uncertainty, and at larger scale even ensemble judges are imperfect.
Source
repowazdogz-droid/evaltrust @ 18e999a5 archived
Repository archived 2 September 2026 after a dated correction to its second experiment. The numbers cited here are from the first experiment and are not affected by the correction, which withdraws a vendor-level interpretation of a separate contrast. Read the correction.
Reproduce
python scripts/build_report.py && python scripts/recompute.py && pytest -q
Independent reproduction
None known.
§ 2

Where this sits

This finding answers Passed is not correct and supports the VERIFY stage of the operating method.

Evidence ledgerBring us a problem