Ω Omega Protocol Bring us a question
MEASURED

Artifact

inspect-replay

Two evaluation logs compared deterministically, distinguishing 'unchanged' from 'cannot tell'.

§ 1

What it establishes

That two Inspect AI evaluation logs can be compared deterministically and sample-aligned, reporting what changed together while distinguishing 'unchanged' from 'cannot tell'.

§ 2

What it does not establish

No causation, co-occurrence only. It does not re-run models; a numeric score change is not labelled a regression; different scorers are not comparable.

§ 3

Method

It reads two recorded evaluation logs, aligns samples, and reports per-sample changes with an explicit ignorance state; the output is deterministic and byte-identical, read-only, and enforced by tests.

§ 4

Results

117 tests pass; the example diff reports CHANGED with the honesty line that no recorded config field accounts for the change.

§ 5

What has to be trusted

The inspect_ai log parser and schema; the author's alignment-key logic, flagged in-repo as the riskiest module; and the UNKNOWN, NOT_CHECKED and NOT_COMPARABLE ignorance model.

§ 6

Prior work

Structured diffing and evaluation reproducibility. The contribution is the deliberate 'unchanged versus cannot tell' distinction with a build-time guard against overclaiming causation.

§ 7

Reproduce it

pip install -e '.[dev]' && inspect-replay compare examples/baseline.eval examples/sample-regression.eval

Findings drawn from this artifact are on the evidence ledger.