← Research index
FRONTIER LAB WATCHReviewed 2 September 2026

Research mathematics needs proof assurance, not a single judge score

OpenAI's correction to a First Proof submission exposes false acceptance as the central metric for AI-assisted mathematics.

Xamit Kadirbekov
Xamit KadirbekovIndependent analysis · Source: OpenAI
Proof verificationEvaluationFormal methods
STATUS · SOURCE REPORT + GERO ANALYSISThis brief has not independently reproduced the laboratory's experiment.
PRIMARY SOURCE · 20 February 2026 Our First Proof submissions ↗

What the laboratory reports

  • OpenAI reports attempts on research-level mathematical problems and distinguishes the sprint from a controlled evaluation.
  • One attempt initially viewed as likely correct was later judged incorrect after expert and community analysis.
  • The report calls for a more rigorous framework for evaluating research-level mathematical work.

The mathematical problem

A proof is a dependency graph, not a scalar score. Its reliability depends on semantic equivalence, hypothesis coverage and the joint false-acceptance rate across necessary steps. Correlated generator and verifier errors invalidate naive independence assumptions.

GERO's proposed response

  • Compile the argument into assumptions, imported results, lemmas, calculations and conclusion dependencies.
  • Route nodes to heterogeneous checkers and reserve expert review for semantic bridges.
  • Publish bounded statuses and replayable evidence rather than an unqualified correct/incorrect label.

A falsifiable experiment

  • Build a sealed set of valid and subtly corrupted research-style proof trajectories.
  • Compare same-family judging, cross-family voting, formal tools and the GERO graph protocol.
  • Use false acceptance of invalid proofs as the primary risk metric.

Read the primary source

This article is an original analytical summary, not a republication. Read Our First Proof submissions for the laboratory's complete claims, methods and context.