Ω Omega Protocol Bring us a question
MEASURED

Artifact

evaltrust

Two off-the-shelf LLM judges fail in opposite directions; a zero-disagreement judge's agreement interval is degenerate.

Archived

Archived 2 September 2026 after a dated correction: the second experiment's vendor-level interpretation was withdrawn; the first experiment's numbers stand and recompute from committed data. Read the correction.

§ 1

What it establishes

That two off-the-shelf language-model judges disagree with the ground truth in opposite directions on a small GSM8K corpus, and that at that sample size a zero-disagreement judge's agreement interval is degenerate, measured, with digests that recompute offline.

§ 2

What it does not establish

Judge reliability. The per-answer correctness labels are assistant-adjudicated, not human; two of the three judges also generated answers, so self-preference is uncorrected; the sample is small.

§ 3

Method

Five graders, three language-model judges, a substring matcher, a numeric oracle, and an ensemble, score 45 real GSM8K trajectories, and every number recomputes from committed data under a digest check.

§ 4

Results

gemma3 disagrees on 4 of 45, all too strict; llama3 on 2 of 45, all too lenient; the ensemble's agreement interval is degenerate at zero disagreements; the report digest is 62bb20c4.

§ 5

What has to be trusted

The objective GSM8K integer answer key; the assistant-adjudicated per-trajectory labels, which are an authored coding scheme and part of the trusted base; the local Ollama judge models at temperature 0.

§ 6

Prior work

Language-model-as-judge evaluation and inter-rater agreement. The contribution is the opposite-direction failure and the degenerate-interval honesty, with every number recomputing from committed data.

§ 7

Reproduce it

python scripts/build_report.py && python scripts/recompute.py && pytest -q

Findings drawn from this artifact are on the evidence ledger.