AI systems can produce verifiable advances on novel mathematical problems rather than merely solve benchmark exercises

Leaning yes
Why — conclusion confidence Moderate: peer-reviewed demonstrations of AI-assisted novel mathematical discovery · machine-checkable constructions and proof artifacts · human involvement remains substantial in problem selection, formalization, and interpretation · limited evidence for autonomous open-ended research-level discovery
Updated 2026-09-15 4 supporting · 3 opposing arguments
PRO 48%CON 52%
Pro 32% · Con 34% — Nuanced 34% — evidence mixed
What the evidence says Evidence quality: High
Graded from the quality of the cited sources · Evidence Protocol

What's this about?

People disagree about whether AI can make new math, not just answer old test questions.

The key question asks if people can check and trust AI’s new ideas.

What supporters say

  • AI has helped math teams spot new links in knot math and other areas, which people then proved.
  • Proof tools such as Lean can check each step that an AI suggests.
  • AI has built hard proofs in set tasks, such as shape puzzles and math contests.
  • AI has found new ways to build things, with computer tests checking each result.

What critics say

  • AI still makes many errors when new problems need many steps of thought.
  • A checked answer may still have little value or meaning for math.
  • Many top AI results still come from set tests or contests, not open research.

How to read this

The number of points on each side does not show who is right; strong proof matters more than a long list.

The bottom line

AI can help find new math results that people or tools can check.

But the best evidence shows AI helps human math work; it does not yet do major research mostly alone.

The fuller picture Reading level: Standard

The claim is partly true: AI systems have helped produce new, checkable mathematical results, but there is little evidence that they can yet carry out important research largely on their own. The strongest record is one of AI-assisted discovery, not independent mathematical research.

The case for

Machine-learning systems have helped mathematicians find previously unknown relationships and conjectures in areas including knot theory and representation theory. Human researchers interpreted those patterns, proved the resulting claims and published them, but the discoveries were not simply answers to standard exercises. This shows that AI can contribute ideas that lead to new, verifiable mathematics.2 (see Figure 1)

There are also more direct examples of machine-generated improvements. FunSearch combined a language model, which proposed computer programs, with automated tests that scored them. It found better constructions for problems such as cap sets and online bin packing. Because each candidate was tested against a clearly defined objective, the improvements could be checked by machines rather than judged only from an AI's explanation.1 (see Figure 2) The limitation is that people designed the representation, the scoring system and the search environment, and the problems were tightly constrained.

Formal proof systems provide another, narrower form of support. AI can propose proof steps that software such as Lean either accepts or rejects. Systems including LeanDojo have helped search existing formal libraries, while other tools have turned informal proof ideas into machine-checked arguments. AlphaGeometry and AlphaProof also produced formal proof traces for difficult geometry and International Mathematical Olympiad problems.3 (see Figure 3) These results demonstrate that AI can construct nontrivial proofs in structured settings, even if they do not show open-ended research ability.4

The case against

Most celebrated results still come from closed, pre-selected tasks. Olympiad problems and theorem-proving benchmarks test difficult reasoning, but they do not require an AI system to decide which questions matter, develop a research program or explain why a result is significant.5 Existing formal libraries also give systems much of the language and structure in advance.

Models remain unreliable when familiar patterns are removed or when a problem requires several unfamiliar steps. Studies report failures in long chains of reasoning, along with plausible but invalid proof steps and difficulty proving new lemmas. That makes it risky to infer broad research ability from strong performance on established benchmarks.6

Most importantly, verification is not the same as importance. A proof assistant can confirm that a conclusion follows from a formal statement. It cannot decide whether the problem is meaningful, whether the formal statement captures the intended mathematics or whether the result changes the field. Formal checking is also limited when an informal conjecture has been mistranslated, incompletely formalized or never successfully proved.7

The bottom line

The evidence favours the claim in a qualified sense, with high confidence that AI can make useful, machine-checkable contributions beyond reproducing benchmark answers. Published discoveries, improved constructions and formal proof artifacts show that these systems can generate new mathematical objects, patterns and proof components.

But the evidence is much weaker for the stronger claim that AI can independently conduct important mathematical research. Humans still usually choose promising problems, design the representations and objectives, interpret the output, judge its importance and complete or validate the proof. The central unanswered question is how much of the apparent novelty comes from the AI's search and synthesis, and how much comes from those substantial human decisions.

Figures & data

Cited sources by side and evidence strengthEach bar counts DISTINCT sources cited on that side, once per source at its highest evidence strength.Supporting3 strong sources32 moderate sources25Opposing2 strong sources21 moderate source13Nuanced4 strong sources44strongmoderate
The evidence base behind this claim: 12 distinct cited sources
Every source cited on this claim, counted once at its highest evidence strength and grouped by the side it supports. Generated from this page's own evidence rows — the same records the verdict is computed from — so the chart and the score cannot disagree. Strength labels follow the scoring methodology.
Davies et al. (2021) figure from Advancing mathematics by guiding human intuition with AI showing the machine-learning workflow that identified a previously unknown connection between knot invariants
The clearest landmark example of AI contributing to genuinely new mathematics: the model detected a non-obvious pattern, while mathematicians interpreted and proved the resulting conjecture. It directly illustrates the distinction between AI-assisted discovery and autonomous proof.
Romera-Paredes et al. (2024) FunSearch figure showing the LLM-plus-evaluator search loop and the quality of newly discovered programs for the cap-set problem and online bin packing, compared with prio
This is the strongest visual example in the evidence of an AI system producing verifiably improved mathematical constructions rather than merely answering benchmark questions. The evaluator establishes objective performance, while the human-designed representation and search setup make the limits of the claim visible.
Trinh et al. (2024) AlphaGeometry performance chart comparing the number of International Mathematical Olympiad geometry problems solved by AlphaGeometry with previous automated systems and human-leve
The iconic benchmark figure for high-level machine-generated mathematical proofs: it demonstrates substantial novel proof construction and formal verification, while the curated olympiad setting helps readers distinguish contest performance from open-ended research discovery.

All contributions are reviewed for clarity, balance, and evidence. The strongest insights are elevated into the argument graph — with credit to you.

Help improve this analysis →
𝕏 Share Facebook LinkedIn