FRONTIER LAB WATCHReviewed 2 September 2026
Research mathematics needs proof assurance, not a single judge score
Permanent archive: Zenodo · 10.5281/zenodo.22729112.
OpenAI's correction to a First Proof submission exposes false acceptance as the central metric for AI-assisted mathematics.
Proof verificationEvaluationFormal methods
PRIMARY SOURCE · 20 February 2026
Our First Proof submissions ↗
What the laboratory reports
- OpenAI reports attempts on research-level mathematical problems and distinguishes the sprint from a controlled evaluation.
- One attempt initially viewed as likely correct was later judged incorrect after expert and community analysis.
- The report calls for a more rigorous framework for evaluating research-level mathematical work.
The mathematical problem
A proof is a dependency graph, not a scalar score. Its reliability depends on semantic equivalence, hypothesis coverage and the joint false-acceptance rate across necessary steps. Correlated generator and verifier errors invalidate naive independence assumptions.
GERO's proposed response
- Compile the argument into assumptions, imported results, lemmas, calculations and conclusion dependencies.
- Route nodes to heterogeneous checkers and reserve expert review for semantic bridges.
- Publish bounded statuses and replayable evidence rather than an unqualified correct/incorrect label.
A falsifiable experiment
- Build a sealed set of valid and subtly corrupted research-style proof trajectories.
- Compare same-family judging, cross-family voting, formal tools and the GERO graph protocol.
- Use false acceptance of invalid proofs as the primary risk metric.
Read the primary source
This article is an original analytical summary, not a republication. Read Our First Proof submissions for the laboratory's complete claims, methods and context.
