Ω Omega Protocol Work together
MEASURED

Finding · authorise

Every language-model agent stayed inside its own allowance. The shared budget was breached in every episode where the allowances added up to more than the budget.

The system said
Every agent within its allowance.
The evidence showed
The shared budget was breached in 60 of 60 BLIND episodes on each of three models, with no individual violation; an informationally redundant restatement of the headroom took it to 4, 2 and 0 of 60.
§ 1

The record

FINDING · commons-compliant-collective-breachauthorise · expressibility
Every language-model agent stayed inside its own allowance. The shared budget was breached in every episode where the allowances added up to more than the budget.
Claim
In a shared-budget environment where each agent's allowance is individually satisfiable but the allowances sum to 1.5 times the collective budget, three language models produced a compliant collective breach in 60 of 60 BLIND episodes each (two population sizes, 30 episodes per cell), with zero individual allowance violations; one appended sentence restating the same headroom took the rate to 4 of 60, 2 of 60 and 0 of 60.
Status
MEASURED Empirical, with a stated sample.
Question
Each compliant is not collectively safe
Subject
Three hosted language models (gpt-5-mini, gpt-4.1-mini, gemini-2.5-flash) acting as agents in a synthetic one-round shared-budget environment (existing models, authored environment). third-party subject
Frame
Collective budget 100; N in {3, 6}; UNSOUND regime cap = ceil(1.5 × 100 / N); 30 episodes per cell fixed before any data; Wilson 95% intervals.
Method
Pre-registered factorial over regime, population size and an information ladder over identical tasks; four amendments each registered before their data; every outcome computed from the environment ledger, never from model text.
Oracle
The environment's own ledger of draws. A stub model whose prose claims it drew nothing while its action draws the full allowance is scored on the action, and that test fails the build if it stops doing so.
Negative control
Present Seven pre-committed controls: an impossible-breach cell that must read 0 of 60 and a forced-breach cell that must read 60 of 60 (both did), an equal-share baseline, an allowance-response check, a benign cell supplying the guardrail's second denominator, a scorer-independence test and five runtime invariant monitors.
Denominator
60 of 60 per model = UNSOUND_N3_BLIND 30 of 30 plus UNSOUND_N6_BLIND 30 of 30. The per-cell table is on the collective-bound page.
Preregistration
Registration commit 3bf280e (2026-08-22) precedes the first model call; amendments 01 to 04 each precede their data; departures in DEVIATIONS.md.
Limitation
A rate here is a fact about a dataset, a prompt and a model on one day, not a capability of a model family or of language-model agents in general. The environment is synthetic and one round deep. Three developers, not four; the guardrail judge shares a vendor with two of the agent arms. No mechanistic claim is made about why the phrasing matters.
Source
repowazdogz-droid/commons-agent-lab @ de3f89d
Reproduce
python3 -m pytest tests -q && python3 -m commons.report _canonical/gpt-5-mini _canonical/gpt-4.1-mini _canonical/gemini-2.5-flash
Independent reproduction
None known.
§ 2

Where this sits

This finding answers Each compliant is not collectively safe and supports the AUTHORISE stage of the operating method. It is an instance of the expressibility mechanism.

Evidence ledgerBring a claim like this one