AI Systems Solve Many, But Not All, Research-Level Math Problems in Rigorous New Benchmark
CAMBRIDGE, MA — First Proof has released the results of its second batch benchmark, assessing the ability of AI systems to autonomously solve naturally occurring mathematical research problems. Four AI systems — OpenAI's ChatGPT 5.5 Pro, and three built on top of commercial AI models by academic teams from ETH Zürich/Aarhus University, UCLA, and Princeton University — were each given ten problems contributed by leading mathematicians, spanning fields from stochastic partial differential equations to combinatorial topology and mathematical logic. Their solutions were evaluated by thirty expert mathematicians in a gathering last week at Harvard’s Center for Mathematical Sciences and Applications.
Each problem arose in the contributor's own research and had a known human solution of at most eight pages that had not previously appeared in print or online. The problems varied in difficulty: some could be solved by a human expert in a few hours, while others took weeks or months.
Across all four systems, a combined total of seven problems received at least one passing grade (rated as essentially flawless or requiring minor revisions). Notably, Problem 5 (stochastic PDE) was solved correctly by one system using a novel approach that differed from the human solution and impressed the referees. Other problems saw complete failure, most notably Problem 4 (metric geometry), where no system made substantial progress. In addition, referees identified two problems where an AI system proposed a potentially viable approach, but significant human effort would be required to repair the submission.
The AI systems tended to perform best when a problem was structurally similar to results already in the literature. A recurring theme in the referee reports was that AI solutions tended to handle routine parts of an argument in meticulous detail while glossing over difficult steps.
"This benchmark gives us insight into what AI can and cannot do in solving well-specified mathematical problems," said the editorial board. "The results show genuine capability. Some solutions were correct, complete, and novel, while others exhibited systematic weaknesses that are useful for the research community to understand. We did not test other key aspects of mathematical research, such as formulating questions or developing mathematical definitions and frameworks."
The benchmark was designed to be independent, transparent, and rigorous. All AI testing was run by First Proof on standardized cloud infrastructure in late May, with each system given 24 hours to solve all ten problems with no human input. The complete record — problems, human solutions, pre-registered author commentary, AI-generated solutions, source code for the three academic systems, referee reports, and system logs including prompts and token costs — is being released publicly today on the First Proof website (1stproof.org).
The First Proof Foundation is a 501(c)(3) nonprofit. This work was supported by the AI For Math Fund, the Survival and Flourishing Fund, and by unrestricted donations from Anthropic and OpenAI which are used towards scientific activities such as problem selection, testing, and grading. The Foundation is grateful to the Simons Institute for the Theory of Computing, which hosted the problem-creation meeting in April, and to Harvard's CMSA, which hosted the grading meeting in June.
---
*Contact:* contact@1stproof.org
*Full report and supplementary materials:* https://1stproof.org