Ω Omega Protocol Bring us a question
MEASURED

Artifact

inspect-audit

A read-only auditor flags silent validity failures in evaluation logs.

§ 1

What it establishes

That a read-only auditor can flag silent validity failures in Inspect evaluation logs, samples dropped from a metric's denominator, grader parse-failures scored as data, a model grading its own output, duplicates, truncation, and non-reproducible settings.

§ 2

What it does not establish

A PASS is not a validity certificate. A biased dataset, a mis-specified task, a subtly-wrong-but-parseable grade, or contamination are out of scope. NOT_CHECKED is not PASS; there are no model calls and no language-model judge.

§ 3

Method

24 static checks run over the log, each firing only on evidence, with catalog-to-code consistency test-enforced; the verdict per check is FAIL, WARN, PASS or NOT_CHECKED.

§ 4

Results

58 tests pass; on the broken fixture, 3 FAIL and 1 WARN over 23 checks (exit 2); on the clean fixture, 0 FAIL (exit 0).

§ 5

What has to be trusted

The author's check catalog and its thresholds, which are an authored coding scheme and part of the trusted base; the inspect_ai schema; and heuristic signals labelled as such.

§ 6

Prior work

Static analysis of evaluation logs and validity threats in measurement. The contribution is the catalog of silent-failure checks with NOT_CHECKED kept distinct from PASS.

§ 7

Reproduce it

pip install -e '.[dev]' && inspect-audit examples/broken.eval

Findings drawn from this artifact are on the evidence ledger.