Ω Omega Protocol Bring us a question

Research

Questions we’ve investigated, and what the evidence showed.

Did it do what it was allowed to do?Approval alone may not show what actually ran.3 findingsExplore

Research distinction: AUTHORISED ≠ EXECUTED

§ 3.1

Authorised is not executed

AUTHORISED ≠ EXECUTED

A policy approves an operation. The operation that runs is not the one approved, and the record of approval reads exactly like a record of what ran.

The question an instrument has to answer. What binds the executed operation to the authorised one, and what would show the binding has failed?

Authorised operation and executed operation are two objects policy: allow 4 KB authorised op ledger: VERIFIED executed op 100,000 bytes written mediation held; the binding between the two operations did not
3 findings
OBSERVED

A program passed all 24 of its recorded tests and still accepted a permission record that had never been issued.

Does not establish. One defect in one program built for this study. Not a measurement of deployed systems, and no claim about how often authorisation records are treated as evidence of authorisation. Not a break-in: producing the forgery requires the ability to run code as the same user, and such an actor can edit the target directly without any record. The minimal fix establishes that a membership check closes this specific defect; it does not establish at-most-once execution, atomic consumption, or independent measurement of the effect.

repowazdogz-droid/omega-n1-permit-membership @ efba311
built for the study
Present
OBSERVED

An agent was limited to 4 KB. It wrote 100 KB. Every test passed, an independent replay verifier said VERIFIED, and a checked invariant held.

Does not establish. One defect in one server built for this study, not a measurement of deployed systems and not a claim about how often authorization layers and executors diverge. Mediation held throughout; what failed was the binding between the authorized operation and the executed one. Policy adequacy is separately not established, and two counterexamples covering it are retained unfixed.

repowazdogz-droid/mcp-authority-boundary @ 03854ddc
built for the study
Present
Artifacts

mcp-authority-boundary · mcp-boundary-audit

Does the record show what actually happened?What the software reports may not match what happened.1 findingExplore

Research distinction: RECORDED ≠ OBSERVED

§ 3.2

Recorded is not observed

RECORDED ≠ OBSERVED

A record says what a component intended. An observation says what changed. A monitor beside the action produces the first and is read as the second.

The question an instrument has to answer. What measures the effect from outside the thing that produced it?

A monitor beside the action and a control on the path action effect monitor beside the path: records "prevented"; the effect already happened action gate effect on the path: a refusal is an effect that did not happen
1 finding
Artifacts

safeguards-control-plane

Can a whole system fail when each part follows its rules?Parts that behave correctly alone can still interact badly.4 findingsExplore

Research distinction: INDIVIDUALLY COMPLIANT ≠ COLLECTIVELY SAFE

§ 3.3

Individually compliant is not collectively safe

INDIVIDUALLY COMPLIANT ≠ COLLECTIVELY SAFE

Every agent stays inside its own limit. The shared quantity they draw on is breached anyway, and no per-agent check can see it.

Interactive exhibit · AMLUCS 2026

When every agent passes and the system still fails

Six agents, one finite pool. Drag a draw, switch the control from per-agent caps to a shared meter, and break each of the three preconditions the aggregate guarantee rests on. A 55-second automatic run walks through the poster's argument.

Read the researchLaunch interactive

The question an instrument has to answer. Which rule can bound the aggregate, under what preconditions, and where does it stop holding?

6 agents, each within its cap of 10. Every local check passes.
6 × 10 = 60 60 > 40: breached
The sum is the one number no local check reads.
4 findings
PROVEN

A spending cap was proved to hold in every state the system can reach, not just the ones anyone tested.

Does not establish. The bound, not conservation. No liveness, availability, or Byzantine model; crash is global in the Lean model, and per-replica crash is only exercised in the bounded checks. The proof holds relative to the transition system being a faithful abstraction of the protocol, which is argued, not machine-checked.

repowazdogz-droid/escrow-budget @ 9c199db4
built for the study
Present
PROVEN

A limit on what an agent may do was proved to hold no matter how concurrent operations interleave.

Does not establish. It is terminal-observation safety, the value the driver reads is within the cap, not an all-intermediate-state invariant, and it carries no liveness or wait-freedom. The kernel check was not re-run in this pass; axiom-freedom rests on the committed assumptions audit plus a live search finding no admitted goals.

repowazdogz-droid/capctl-iris @ 5e9284a0
built for the study
Present
MEASURED

Blind to the pool, the budget was breached in 60 of 60 episodes on each of three models with no agent over its own allowance. With the pool state shown the budget was still breached on every model, at rates that differ by model; the per-cell counts, including episodes where an agent exceeded its own allowance, are in the table on the collective-bound page.

Does not establish. A rate here is a fact about a dataset, a prompt and a model on one day, not a capability of a model family or of language-model agents in general. The environment is synthetic and one round deep. Three developers, not four; the guardrail judge shares a vendor with two of the agent arms. No mechanistic claim is made about why the phrasing matters.

repowazdogz-droid/commons-agent-lab @ de3f89d
third-party subject
Present
PROVEN

Six agents each inside a cap of 10 breach a pool of 40; a single rule on the sum removes every such breach, while a harm equal to one agent's draw (threshold 15, below the pool) escapes it and a per-recipient cap closes it.

Does not establish. Anything about a deployed multi-agent system. The Lean results are over closed data; the Z3 results are verdicts over an encoding whose faithfulness to the informal model is not machine-checked. The per-agent cap is not shown to be unfixable: a local allowance of floor(G/N) is collectively safe when population and utilisation are known and fixed. The channel-blindness result takes perfectly secure steganography as an assumption, not a construction. Not preregistered.

repowazdogz-droid/collective-bound @ public commit of 2026-09-02
built for the study
Present
Artifacts

capctl-iris · escrow-budget · commons-agent-lab · collective-bound

Can a passing test still miss a problem?A pass only tells you what the test checked.13 findingsExplore

Research distinction: PASSED ≠ CORRECT

§ 3.4

Passed is not correct

PASSED ≠ CORRECT

A judge, a test suite or a verifier says pass. The property it stands for has failed, and the pass carries no information about it.

The question an instrument has to answer. What does this checker actually distinguish, and against which oracle?

13 findings
OBSERVED

The same failing evaluation reported a different number of finished samples when its scoring steps were listed in a different order.

Does not establish. One defect in one framework, in the metadata a finished run records about itself. It says nothing about the correctness of any model score, nothing about how often the condition arises in practice, and nothing about other evaluation frameworks. The upstream design decision to score errored samples is separate and was not disputed.

UKGovernmentBEIS/inspect_ai @ f0e57a7c
third-party subject
Not applicable
MEASURED

Two AI graders agreeing tells you less than it appears to, because part of that agreement is forced by both being right.

Does not establish. Not a claim that agreement between judges carries no information, and not a general law about language-model judges. It is a property of these judge checkpoints on these 600 items against a machine-checkable answer key. Ground truth here is a key, not human labelling, and the mechanism behind the second experiment is reported as open rather than settled.

repowazdogz-droid/evaltrust @ c8196e1a archived
third-party subject
Present
MEASURED

A language-model judge scored 8 out of 10 on decisions that a solver proved violated the encoded policy.

Does not establish. The judge is qwen2.5-coder:14b, a local model, not a frontier judge, and on model-authored rows it grades its own output. A stronger judge may do better; that is unmeasured. The harness depends on two packages that are not yet published on their own, vendored with provenance. Six cases is a demonstration of the split, not a rate.

repowazdogz-droid/proof-carrying-evals @ f22def1
built for the study
Present
OBSERVED

A local language-model worker emptied a 415-line module to two lines. The repository’s 24 tests still passed. An independent syntax-tree gate, in series before any merge, rejected the patch.

Does not establish. One worker model on ten files from the author’s own repositories on one day; no rate for any model family and nothing about deployed systems. The gate certifies syntactic equivalence, not correctness: removing an unused import that had an import-time side effect would pass it. The lint checker’s own reading of the truncated file was recorded green in one arm and red in the other on the identical patch, so the checker verdict is not relied on here; the test result is. The record is not public, and independent reproduction is none.

artifact not public
third-party subject
Present
OBSERVED

An error-correcting decoder in an open hardware library declared a syndrome output that nothing drove. A cover requirement added while proving the encoder and decoder pair found it; the one-line fix and the proof were reviewed and merged by the maintainer.

Does not establish. One undriven output in one module of one library. The proof covers the pair at widths 1, 2, 4, 5, 11 and 12 by default and 128 properties up to width 64 in a sweep; it says nothing about the data output under a double error, nothing beyond two flipped bits, and nothing about other modules. The proof needs SymbiYosys, Yosys and abc; it was run by the author before and after the fix and reported in the pull request, and has not been re-run for this record. Maintainer acceptance is review of one change, not endorsement of Omega. The fix is on the default branch and in no tagged release as of 20 September 2026.

pulp-platform/common_cells @ a257b714
third-party subject
Present
OBSERVED

A battery model ran to completion on every revision tested and lost 2.4% of its lithium. Independent inventory accounting localised the loss to one equation, and a project contributor confirmed the missing term.

Does not establish. One model in one library under one protocol and one parameter set. Nothing against PyBaMM's full DFN, whose control conserved to 3.5e-14 mol, and nothing about the finite-volume discretisation of nonlinear source terms, which a contributor raised and which was not investigated. Eleven of 84 half-cell jobs failed inside the solver at tight tolerance, so no half-cell spatial convergence order is claimed. No physical battery was measured. The #5745 fix is on the default branch but not in a tagged release; #5746 remains open.

repowazdogz-droid/pybamm-5700-lithium-inventory @ adaa877
third-party subject
Present
MEASURED

A proposed analytic gradient for a quantum-chemistry method matched independently computed finite differences in 22 of 22 cells; the released code path failed in 22 of 22. One of the referee's own thresholds fired on both arms alike and was reported as an instrument defect.

Does not establish. Agreement on 11 molecules establishes agreement on those 11 molecules. Not tested: open-shell references, periodic systems, symmetry-adapted molecules, systems larger than six atoms, auxiliary bases beyond the two used, or the geometry-optimiser path. The finite-difference referee differentiates the same energy code the pull request differentiates, so it establishes consistency with that code, not agreement with an external truth. The origin of the shared 1e-7 rotational residual is not identified. The evidence has not been posted to the pull request; this is an internal Omega record.

artifact not public
third-party subject
Present
OBSERVED

A compiled PyTorch function returned a result computed for a different object state, with no recompilation, where the uncompiled function returned the right one. The investigation was carried out by a locally hosted model with nothing leaving the machine.

Does not establish. One issue, one pinned build, one run. Nothing about the inductor backend, unions, method-bearing protocols, larger programs or the second scored call order. The upstream issue was already open and is not Omega's discovery; nothing was posted to it and no fix was proposed. The worker's own directional claim about which call order recompiles is unscored, conflicts with the hidden matrix, and is not evidence; its empty list of untested dimensions overstates completeness. No root cause was verified in the compiler's source. The result is a fact about this task, not a general capability of the local model.

artifact not public
third-party subject
Present
OBSERVED

A convincing-looking Unreal validator was rejected because it reconstructed transforms from the source CSV instead of observing instantiated geometry. Three planted geometry corruptions left its result byte-identical.

Does not establish. This does not show that the original Unreal scene was geometrically wrong, and no repaired Unreal validator was built or run against the scene. The regression fixture reproduces the data dependency locally; it is not an independent reproduction in Unreal. Instrument Qualifier v0 used seven authored dependency graphs, so its recall on real validators is unknown and no general qualification capability is claimed.

artifact not public
built for the study
Present
OBSERVED

A three-phase power-flow solver reported convergence for four transformer types it does not support, returning bus voltages near a million per unit with no warning. Ten other unsupported types were correctly rejected.

Does not establish. One network, one load, zero phase shift, one commit. The four silent strings map to one model, so they are one observation, not four. No independent oracle was used for the supported groups, so plausible is the ceiling of what is said about them. Whether the zero-sequence skip is intended for another calculation path is upstream's call. The issue is open and unanswered; nothing has been fixed.

artifact not public
third-party subject
Present
Artifacts

evaltrust · proof-carrying-evals · evidence-audit

Does the proof answer the question that matters?A proof can be correct while covering a narrower question.2 findingsExplore

Research distinction: SPECIFICATION PROVED ≠ PROPERTY INTENDED

§ 3.5

The specification proved is not the property intended

SPECIFICATION PROVED ≠ PROPERTY INTENDED

Every property written from the specification holds. The requirement the design was meant to satisfy was never written down, so nothing checks it.

The question an instrument has to answer. Which requirements have anything checking them, and which formalism can even state them?

Properties derived from a specification cover part of the intended requirement the requirement the design was meant to satisfy properties from the spec all pass never stated nothing checks it two injected defects passed every property
2 findings
OBSERVED

Two injected hardware defects passed every property written from the original specification, because the specification never stated the requirement they break.

Does not establish. A claim about this ~450-line design and this property set only. The five mutations were written by the same person who wrote the properties; a 200-mutant Yosys run was added for that reason and showed the hand-written mutations had probed the wrong part of the design. Nothing here is a statement about any commercial verification flow.

repowazdogz-droid/spcu-verification @ e49a0c5
built for the study
Present
Artifacts

vsf-cjson · spcu-verification

Can someone else check the result for themselves?Showing the steps is different from repeating the check independently.2 findingsExplore

Research distinction: REPLAYABLE ≠ INDEPENDENTLY ESTABLISHED

§ 3.6

Replayable is not independently established

REPLAYABLE ≠ INDEPENDENTLY ESTABLISHED

A record is intact, its integrity rechecks, and its author can re-derive the result on the author’s machine. Whether anyone else, on other hardware, can establish it is a different question with a different answer.

The question an instrument has to answer. Can the result be re-derived from the record alone, by someone else, and where does re-derivation stop?

2 findings
Artifacts

nanogpt-provenance · OMEGA

Accepted changes

12 fixes or contributions accepted · 4 proposed by Omega

Checked 7 October 2026. The others were made by maintainers or contributors. A merged change means the project accepted the change, not that it endorses Omega. See the counting rules and sources.

See all accepted changes →
More on Omega’s researchExplore
§ 3

Research

Omega studies how to settle important technical questions with evidence others can check, and where a check claims more than it actually observed. This page keeps the questions, instruments, controls, results and limits in full. The work is organised by which distinction came apart, because that guides the next experiment.

Open questions are part of the programme. Research groups and domain experts can bring a question, a competing explanation or an experimental setting for a joint investigation. See the research partnership path or start a conversation.

6 questions25 findings in this site ledger19 artifacts13 on third-party systems12 accepted upstream