Ω Omega Protocol Work together
MEASURED

Artifact

commons-agent-lab

A pre-registered study of whether LLM agents produce the collective failure that per-agent rules permit.

§ 1

What it establishes

That three hosted language models, acting as agents in a shared-budget environment, produce a compliant collective breach in every BLIND episode where the allowances exceed the budget; that an aggregate-channel guardrail detects every pivotal action while a local-channel guardrail is analytically blind; and that one redundant sentence about the headroom moves the breach rate more than the information it carries.

§ 2

What it does not establish

Anything about a deployment or a product. A rate is a fact about a dataset, a prompt and a model on one day. Three developers, not four. No mechanistic claim about why phrasing matters.

§ 3

Method

One-round shared-budget commons; factorial over regime, population and an information ladder; four amendments each registered before their data; seven pre-committed controls.

§ 4

Results

Compliant breach 60 of 60 in BLIND on each model; PEERSUM 0, 7 and 2 of 60; RESTATE 4, 2 and 0 of 60; aggregate-channel LLM judge 100 of 100 detections with 0 of 90 false blocks and 0 breaches in series; impossible-breach 0 of 60 and forced-breach 60 of 60 controls fired.

§ 5

What has to be trusted

The environment and scorer (authored, with a scorer-independence test); the pre-registered design fingerprint; hosted model APIs on the run dates; the guardrail judge gpt-5-mini, which shares a vendor with two agent arms.

§ 6

Prior work

Commons and public-goods games; per-agent guardrails as the common deployment shape. The contribution is the pre-registered empirical counterpart to a formal result, with mechanical scoring and controls that fire.

§ 7

Reproduce it

python3 -m pytest tests -q && python3 -m commons.report _canonical/gpt-5-mini _canonical/gpt-4.1-mini _canonical/gemini-2.5-flash

Findings drawn from this artifact are on the evidence ledger.