The celebrated breakthroughs by AI in solving long-standing mathematical problems are primarily counterexamples rather than comprehensive proofs, according to Fields Medalist Timothy Gowers. While recognizing the impressive progress of large language models, Gowers underscores the nuanced difference between these two types of mathematical results.
- Counterexamples invalidate conjectures with single instances; proofs require exhaustive arguments.
- AI models have excelled at producing counterexamples for difficult math problems.
- Human-like judgment remains crucial for identifying promising proof strategies.
What happened
OpenAI recently announced the solving of ten open problems in mathematics and theoretical computer science, including the construction of a non-sofic group and establishing superexponential growth of a multicolour Ramsey number. These claims garnered broad attention as significant AI-driven mathematical achievements. However, Timothy Gowers, a Fields Medal-winning mathematician, reviewed these results and offered a detailed perspective on their actual nature.
Gowers confirmed that while these breakthroughs are extraordinarily impressive, the majority are counterexamples rather than universal proofs. He explains that a counterexample disproves conjectures by providing a single case against them, whereas a proof universally validates a statement. Importantly, Gowers observes that AI’s strongest achievements to date have predominantly been in generating counterexamples, which leverages the AI’s capacity to test many possibilities rapidly.
Why it matters
This distinction matters because it illuminates the different kinds of mathematical work and the challenges AI faces in replicating them. While AI can rapidly explore vast mathematical objects to find counterexamples, proving that a statement holds universally demands understanding, insight, and nuanced judgment that go beyond raw computational power.
Gowers highlights that human mathematicians use intuition and experience to decide which approaches to continue pursuing and which to abandon. This selective process, or having a “nose” for promising directions, allows humans to prune overwhelmingly large search spaces effectively — a skill AI currently lacks. Thus, despite AI’s remarkable progress, collaborative interplay between human insight and machine speed remains essential.
What to watch next
The evolution of AI in mathematics will hinge on addressing this gap in judgment and strategic reasoning. Improvements in AI’s ability to evaluate the worthiness of incomplete proofs and selectively refine them could enable it to produce more complete and universally valid proofs in the future.
Researchers and mathematicians will be closely monitoring forthcoming AI developments, such as enhancements in large language models like GPT-5.6 Pro, to see if these systems begin to demonstrate better 'nose' and strategic decision-making. Additionally, further collaborations between AI and human experts may unlock new layers of mathematical discovery that neither could achieve alone.