§ 7
Findings
The findings live on the evidence ledger, with every field.
- A program passed all 24 of its recorded tests and still accepted a permission record that had never been issued.
- An agent was limited to 4 KB. It wrote 100 KB. Every test passed, an independent replay verifier said VERIFIED, and a checked invariant held.
- A spending cap was proved to hold in every state the system can reach, not just the ones anyone tested.
- A widely used library reads a number, writes it back out, and the two are not the same.
- The same failing evaluation reported a different number of finished samples when its scoring steps were listed in a different order.
- A limit on what an agent may do was proved to hold no matter how concurrent operations interleave.
- Two AI graders agreeing tells you less than it appears to, because part of that agreement is forced by both being right.
- At the sample size used, the question the experiment was built to answer could not be answered either way, and that is reported rather than rounded to a result.
- Two off-the-shelf AI graders scoring the same work were wrong in opposite directions.
- A dashboard reported 1,026 actions prevented. All 1,026 had already executed.
- Changing a sealed record after the fact is provably as hard as breaking the hash it is sealed with.
- Tools removed from the list an agent is shown were still there to call by name.
- A training run re-derives its exact result from its recorded inputs on one machine. Across different hardware it stays unknown.
- Blind to the pool, the budget was breached in 60 of 60 episodes on each of three models with no agent over its own allowance. With the pool state shown the budget was still breached on every model, at rates that differ by model; the per-cell counts, including episodes where an agent exceeded its own allowance, are in the table on the collective-bound page.
- A language-model judge scored 8 out of 10 on decisions that a solver proved violated the encoded policy.
- Two injected hardware defects passed every property written from the original specification, because the specification never stated the requirement they break.
- Three green verification runs could not have failed: an unreached assertion, a model checker that instrumented nothing, and an axiom audit identical for correct and wrong models.
- Six agents each inside a cap of 10 breach a pool of 40; a single rule on the sum removes every such breach, while a harm equal to one agent's draw (threshold 15, below the pool) escapes it and a per-recipient cap closes it.
- A local language-model worker emptied a 415-line module to two lines. The repository’s 24 tests still passed. An independent syntax-tree gate, in series before any merge, rejected the patch.
- An error-correcting decoder in an open hardware library declared a syndrome output that nothing drove. A cover requirement added while proving the encoder and decoder pair found it; the one-line fix and the proof were reviewed and merged by the maintainer.
- A battery model ran to completion on every revision tested and lost 2.4% of its lithium. Independent inventory accounting localised the loss to one equation, and a project contributor confirmed the missing term.
- A proposed analytic gradient for a quantum-chemistry method matched independently computed finite differences in 22 of 22 cells; the released code path failed in 22 of 22. One of the referee's own thresholds fired on both arms alike and was reported as an instrument defect.
- A compiled PyTorch function returned a result computed for a different object state, with no recompilation, where the uncompiled function returned the right one. The investigation was carried out by a locally hosted model with nothing leaving the machine.
- A convincing-looking Unreal validator was rejected because it reconstructed transforms from the source CSV instead of observing instantiated geometry. Three planted geometry corruptions left its result byte-identical.
- A three-phase power-flow solver reported convergence for four transformer types it does not support, returning bus voltages near a million per unit with no warning. Ten other unsupported types were correctly rejected.