MEASURED
Finding · replay
A training run re-derives its exact result from its recorded inputs on one machine. Across different hardware it stays unknown.
- The system said
- Replay: VERIFIED, bit for bit.
- The evidence showed
- On one machine, with 12 of 12 controls firing. On other hardware the answer is untested and is recorded as UNKNOWN rather than assumed. Independent reproduction by anyone else: none.
§ 1
The record
FINDING · training-run-rederives-bit-for-bitreplay
A training run re-derives its exact result from its recorded inputs on one machine. Across different hardware it stays unknown.
- Claim
- A training run's final loss and trajectory re-derive bit-for-bit from its recorded inputs on one machine; cross-hardware re-derivation is left as UNKNOWN, and not claimed.
- Status
- MEASURED Empirical, with a stated sample.
- Subject
- A small training run and its verifier, on one machine (authored). built for the study
- Frame
- Python 3.14.4, NumPy 2.4.4 on Apple Accelerate; cross-hardware untested.
- Method
- The verifier re-runs training from the recorded configuration and compares the trajectory and final loss bit-for-bit.
- Oracle
- The recorded final loss and trajectory fingerprint, recomputed rather than rechecked.
- Negative control
- Present A 12-case suite triggers each verdict state (VERIFIED, TAMPERED, UNKNOWN) with a real tamper.
- Denominator
- 12 of 12 controls; one machine.
- Limitation
- Bit-for-bit reproducibility on other hardware, which is untested and reported UNKNOWN rather than upgraded. It does not establish code correctness, buggy-but-faithful code still verifies, nor result quality nor accountability.
- Source
- repowazdogz-droid/nanogpt-provenance @ 732da8b1
- Reproduce
./run_all.sh
- Independent reproduction
- None known.
§ 2
Where this sits
This finding answers Replayable is not independently established and supports the REPLAY stage of the operating method.