Upstream record
Problems found in other people’s software, and what happened next. Solver, simulation and hardware items come first; the rest are bugs in AI-evaluation and developer tooling. A merged fix is a maintainer’s review of one change, not an endorsement of me. Twelve led to merged changes: I wrote four of them, and the projects’ own developers wrote the rest after my reports (one of those is an opt-in mitigation rather than a change of default). One is still open. All of this was self-initiated. Every row links to the public fix or issue; status was checked on 8 October 2026.
| Project | What was wrong | My part | Status |
|---|---|---|---|
| cvc5 | SMT solver returned a wrong “sat” on a separation-logic problem (reported by another user). | Root cause and fix | Released in 1.4.2 2026-10-01 |
| PyBaMM | BasicDFN with variable transference lost 2.44% of its lithium. | Measured, diagnosed and fixed (first reported in #5700) | Released in 26.9.0.0 My report 2026-09-18 |
| PyBaMM | Half-cell lithium diagnostic used the wrong length scale (factor 11.45). | Reported; fixed by a maintainer | Released in 26.9.0.0 My report 2026-09-22 |
| PULP common_cells | ECC decoder output declared and never driven; first formal proof of the module. | Found and fixed | Merged 2026-09-03 |
| pandapower | Four unsupported transformer types returned converged=True with impossible results. | Reported | Open 2026-09-06 |
| inspect-robots | Run records did not include the grader configuration, so differently graded runs looked identical. | Found and fixed | Merged 2026-09-21 |
| inspect-robots | Logs did not record which path produced each operator judgement. | Reported; fixed by a maintainer | Merged My report 2026-08-31 |
| Inspect AI (UK AI Security Institute) | The same failing run reported 6 of 6 or 4 of 6 completed samples depending on scorer order. | Reported; fixed by a contributor | Merged My report 2026-07-29 |
| LemmaScript | “1 verified, 0 errors” on code that called a random function, modelled as pure. | Reported; the maintainer added an opt-in annotation (default unchanged) | Mitigated in v0.6.1 My report 2026-08-25 |
| con-leche (Lean) | The checker command could exit successfully without running the checker. | Reported; fixed by a maintainer | Merged My report 2026-09-10 |
| NVIDIA labs-OO-Agents | Generator methods were wrapped as coroutines, so tracing spans ended before any work ran. | Reported; fixed by a maintainer | Merged My report 2026-08-31 |
| NVIDIA SkillSpector | Every scan through one model provider failed on a schema the endpoint rejected. | Reported; resolved in a maintainer’s release sync | Merged My report 2026-06-16 |
| Orbit | A documented clean install failed at import: a dependency was undeclared. | Reported; fixed by a maintainer | Merged My report 2026-09-17 |
About this record
Methods used on the public record: independent recomputation of physical quantities, differential runs, pre-registered test matrices, formal proofs (SymbiYosys, Lean, Z3, TLA+), and root-causing in C++ and Python.
My longer research notes, including formal proofs and AI-evaluation studies, are published under the name Omega: 46 public results with sources, 40 with re-run instructions. Omega is my notebook, not a company. For open-ended scientific or engineering R&D rather than one result, see my R&D practice.