Warren SmithIndependent checks of technical results Ask for a check

How Claim Check works

Send one technical claim your team is about to rely on. I build an independent check, show that it catches a planted failure, run it on your pinned version, and report what holds, what fails and what can’t be told yet.

What do I send?

To start, one sentence: the claim, and the decision that depends on it. Later, the thing itself: a repository and commit, a notebook, a model and its inputs, a results file, or access to a system.

Claims that fit well:

  • “Our solver conserves mass on these inputs.”
  • “The GPU path gives the same answer as the CPU path.”
  • “This proof covers the decoder’s outputs.”
  • “This eval score measures refusals correctly.”
  • “This result in our paper reproduces from the released code.”

What happens?

  1. Scope. I reply with whether the claim can be tested with what’s available, what access I’d need, how long it will take and a fixed fee. If it can’t be tested, I explain why. That reply costs nothing.
  2. Agreement. The wording of the claim, the version, inputs, scope, duration and fee are agreed in writing. An NDA is fine.
  3. Early readout. Once access is in place: does the result reproduce at all, and has anything already broken?
  4. The check. Built independently of the check that passed. Before I trust a pass, I show the check fails on a known-bad case.
  5. Report and walkthrough. A written report in the structure of the sample, and a call to go through it.

What do you test?

Whether the claim, as agreed, holds on the version and inputs agreed. Usually some mix of:

  • recomputing a quantity from raw outputs rather than the tool’s own totals;
  • running the same input two ways;
  • pushing inputs to the edge of what’s supported;
  • planting a fault to see whether your existing check notices;
  • checking whether a proof or test actually reads the output you care about.

What do I receive?

  • A verdict: Holds, Fails or Can’t tell yet, scoped to the version and inputs.
  • The evidence: what ran, on what, with what result.
  • Proof the check works: the planted failure it caught.
  • The limits: what the result does not show.
  • A script to re-run the check on your next release, where the method allows.
  • The smallest useful next step, if there is one.

See a sample report, built from a real public case in exactly this structure.

How long might it take?

I expect most single claims to take one to two weeks of elapsed time, with the early readout first. The duration is agreed before work starts. This is an estimate based on the public investigations.

What if the answer is negative?

Fails: you hear it privately, before your decision, with the smallest reproducer and a suggested fix. Can’t tell yet: you get the reason (missing access, no independent reference, behaviour outside the agreed scope) and what would settle it. The fee is the same either way.

What stays private?

  • Everything you send, and the result, is confidential by default. Nothing is published, and the work is not mentioned, without your written agreement.
  • Sensitive data can stay in your environment. I can supply a check for your team to run there and review only its outputs.
  • I don’t change a result to suit anyone. You can add context, and it will be included.
  • If the check finds a bug in a third-party open-source dependency, we agree together how and when to report it upstream.

What’s the next step after a report?

A fix and a re-check, turning the check into a test that runs on every release, a second claim, or nothing.

What it isn’t

Not certification, accreditation or a regulated audit. Not a penetration test. Not a guarantee about anything beyond the claim, version and inputs checked. I hold no accreditation or security clearance.

Ask for a check

What’s going on?

What I would do first

Find out whether the check can fail at all. I plant a fault the check should catch and see whether it notices, then look at what the passing check actually reads.

What to send

The check or test, the commit, and one passing run.

Closest real case

PULP common_cells: My first formal proof of a decoder went green while one of its outputs was never connected.

Ask about this

What I would do first

Recompute a quantity that has to balance (mass, charge, energy, money) from the raw outputs, not from the tool’s own totals, and see where it stops balancing.

What to send

The model, its inputs, and the run that looks wrong.

Closest real case

PyBaMM: A battery model solved cleanly and lost 2.44% of its lithium.

Ask about this

What I would do first

Pin every version and input, re-run, then narrow the difference to code, data, environment or hardware, one change at a time.

What to send

The original result, the code and data you have, and what you got instead.

Closest real case

PyBaMM: Three revisions frozen and hashed; the same loss on every one.

Ask about this

What I would do first

Restate the claim so that it could fail, agree in advance what would settle it, then test only that, and report what it doesn’t cover.

What to send

The claim, the decision it feeds, and the date.

Closest real case

pandapower: 18 transformer types, each result tested against five physical checks.

Ask about this

What I would do first

Push inputs to the edges of what is supported, run the same input two ways, and look for a change that should do nothing but alters the answer.

What to send

Access to the system or code, and what it must never do.

Closest real case

cvc5: An always-true line changed a solver’s answer.

Ask about this

Not a fit: penetration testing, certification or regulated audits, or building a product. I’ll say so in the first reply and suggest who to ask if I can.

Why not just…

…ask our own engineers?
They’re the right people to fix it. An outside check is useful because it doesn’t share the assumptions the original was built on.
…ask an AI assistant?
Fine for a first opinion. It doesn’t run your pinned version, plant a fault to test its own check, or stand behind the result.
…hire a consultancy?
Right for a programme. For one claim before a deadline, the scoping alone often costs more than the check.
…find an academic collaborator?
Right for open research questions. Claim Check works to your decision date, not a publication cycle.

Ask for a check

You don’t need to know how to test it. I reply with a testable version of your question, what I would need, how long it would take and a fixed fee. If it can’t be tested, I say why.

This page sends nothing anywhere. The button opens your own email app with the message below, and you can read every word before you send it. Please keep confidential material out of a first email; an NDA is fine once we know it’s a fit.

Or write directly: warren@omegaprotocol.org

The email that will open
To: warren@omegaprotocol.org
Subject: Claim Check