How Claim Check works
Send one technical claim your team is about to rely on. I build an independent check, show that it catches a planted failure, run it on your pinned version, and report what holds, what fails and what can’t be told yet.
What do I send?
To start, one sentence: the claim, and the decision that depends on it. Later, the thing itself: a repository and commit, a notebook, a model and its inputs, a results file, or access to a system.
Claims that fit well:
- “Our solver conserves mass on these inputs.”
- “The GPU path gives the same answer as the CPU path.”
- “This proof covers the decoder’s outputs.”
- “This eval score measures refusals correctly.”
- “This result in our paper reproduces from the released code.”
What happens?
- Scope. I reply with whether the claim can be tested with what’s available, what access I’d need, how long it will take and a fixed fee. If it can’t be tested, I explain why. That reply costs nothing.
- Agreement. The wording of the claim, the version, inputs, scope, duration and fee are agreed in writing. An NDA is fine.
- Early readout. Once access is in place: does the result reproduce at all, and has anything already broken?
- The check. Built independently of the check that passed. Before I trust a pass, I show the check fails on a known-bad case.
- Report and walkthrough. A written report in the structure of the sample, and a call to go through it.
What do you test?
Whether the claim, as agreed, holds on the version and inputs agreed. Usually some mix of:
- recomputing a quantity from raw outputs rather than the tool’s own totals;
- running the same input two ways;
- pushing inputs to the edge of what’s supported;
- planting a fault to see whether your existing check notices;
- checking whether a proof or test actually reads the output you care about.
What do I receive?
- A verdict: Holds, Fails or Can’t tell yet, scoped to the version and inputs.
- The evidence: what ran, on what, with what result.
- Proof the check works: the planted failure it caught.
- The limits: what the result does not show.
- A script to re-run the check on your next release, where the method allows.
- The smallest useful next step, if there is one.
See a sample report, built from a real public case in exactly this structure.
How long might it take?
I expect most single claims to take one to two weeks of elapsed time, with the early readout first. The duration is agreed before work starts. This is an estimate based on the public investigations.
What if the answer is negative?
Fails: you hear it privately, before your decision, with the smallest reproducer and a suggested fix. Can’t tell yet: you get the reason (missing access, no independent reference, behaviour outside the agreed scope) and what would settle it. The fee is the same either way.
What stays private?
- Everything you send, and the result, is confidential by default. Nothing is published, and the work is not mentioned, without your written agreement.
- Sensitive data can stay in your environment. I can supply a check for your team to run there and review only its outputs.
- I don’t change a result to suit anyone. You can add context, and it will be included.
- If the check finds a bug in a third-party open-source dependency, we agree together how and when to report it upstream.
What’s the next step after a report?
A fix and a re-check, turning the check into a test that runs on every release, a second claim, or nothing.
What it isn’t
Not certification, accreditation or a regulated audit. Not a penetration test. Not a guarantee about anything beyond the claim, version and inputs checked. I hold no accreditation or security clearance.