Artifact
inspect-replay
Two evaluation logs compared deterministically, distinguishing 'unchanged' from 'cannot tell'.
What it establishes
That two Inspect AI evaluation logs can be compared deterministically and sample-aligned, reporting what changed together while distinguishing 'unchanged' from 'cannot tell'.
What it does not establish
No causation, co-occurrence only. It does not re-run models; a numeric score change is not labelled a regression; different scorers are not comparable.
Method
It reads two recorded evaluation logs, aligns samples, and reports per-sample changes with an explicit ignorance state; the output is deterministic and byte-identical, read-only, and enforced by tests.
Results
117 tests pass; the example diff reports CHANGED with the honesty line that no recorded config field accounts for the change.
What has to be trusted
The inspect_ai log parser and schema; the author's alignment-key logic, flagged in-repo as the riskiest module; and the UNKNOWN, NOT_CHECKED and NOT_COMPARABLE ignorance model.
Prior work
Structured diffing and evaluation reproducibility. The contribution is the deliberate 'unchanged versus cannot tell' distinction with a build-time guard against overclaiming causation.
Reproduce it
pip install -e '.[dev]' && inspect-replay compare examples/baseline.eval examples/sample-regression.eval
Findings drawn from this artifact are on the evidence ledger.