← Research index
FRONTIER LAB WATCHReviewed 2 September 2026
Research mathematics needs proof assurance, not a single judge score
OpenAI's correction to a First Proof submission exposes false acceptance as the central metric for AI-assisted mathematics.
Xamit KadirbekovIndependent analysis · Source: OpenAI
Proof verificationEvaluationFormal methods
STATUS · SOURCE REPORT + GERO ANALYSISThis brief has not independently reproduced the laboratory's experiment.
PRIMARY SOURCE · 20 February 2026
Our First Proof submissions ↗
What the laboratory reports
OpenAI reports attempts on research-level mathematical problems and distinguishes the sprint from a controlled evaluation.
One attempt initially viewed as likely correct was later judged incorrect after expert and community analysis.
The report calls for a more rigorous framework for evaluating research-level mathematical work.
The mathematical problem
A proof is a dependency graph, not a scalar score. Its reliability depends on semantic equivalence, hypothesis coverage and the joint false-acceptance rate across necessary steps. Correlated generator and verifier errors invalidate naive independence assumptions.
GERO's proposed response
Compile the argument into assumptions, imported results, lemmas, calculations and conclusion dependencies.
Route nodes to heterogeneous checkers and reserve expert review for semantic bridges.
Publish bounded statuses and replayable evidence rather than an unqualified correct/incorrect label.
A falsifiable experiment
Build a sealed set of valid and subtly corrupted research-style proof trajectories.
Compare same-family judging, cross-family voting, formal tools and the GERO graph protocol.
Use false acceptance of invalid proofs as the primary risk metric.
Boundary
A formal certificate proves the encoded statement. Separate review is still required to establish that the encoding matches the intended natural-language claim.
Read the primary source
This article is an original analytical summary, not a republication. Read Our First Proof submissions for the laboratory's complete claims, methods and context.
