Artifact
inspect-audit
A read-only auditor flags silent validity failures in evaluation logs.
What it establishes
That a read-only auditor can flag silent validity failures in Inspect evaluation logs, samples dropped from a metric's denominator, grader parse-failures scored as data, a model grading its own output, duplicates, truncation, and non-reproducible settings.
What it does not establish
A PASS is not a validity certificate. A biased dataset, a mis-specified task, a subtly-wrong-but-parseable grade, or contamination are out of scope. NOT_CHECKED is not PASS; there are no model calls and no language-model judge.
Method
24 static checks run over the log, each firing only on evidence, with catalog-to-code consistency test-enforced; the verdict per check is FAIL, WARN, PASS or NOT_CHECKED.
Results
58 tests pass; on the broken fixture, 3 FAIL and 1 WARN over 23 checks (exit 2); on the clean fixture, 0 FAIL (exit 0).
What has to be trusted
The author's check catalog and its thresholds, which are an authored coding scheme and part of the trusted base; the inspect_ai schema; and heuristic signals labelled as such.
Prior work
Static analysis of evaluation logs and validity threats in measurement. The contribution is the catalog of silent-failure checks with NOT_CHECKED kept distinct from PASS.
Reproduce it
pip install -e '.[dev]' && inspect-audit examples/broken.eval
Findings drawn from this artifact are on the evidence ledger.