2026-09-21AITao

Terence Tao: SAIR's Open Math Model and Proof Indigestion

In a keynote at Caltech, Terence Tao analyzes how fast AI proofs cause proof indigestion without human explanation, and introduces SAIR's open math model initiative and competition experiments.

Contents7 sections
  1. Beyond Problem Solving: The Complete Mathematical Lifecycle
  2. "Proof Indigestion" and the Understanding Bottleneck
  3. 25 Fields Medalists on AI Misalignment: When Grading Replaces Understanding
  4. SAIR’s Open Math Model: Returning Tools to Researchers
  5. The Distillation Challenge: How a 10KB Prompt Boosted Open Models by 30% to 40%
  6. The Inverse Galois Challenge: From Competition to a Collaborative Super-Team
  7. Data Pollution and Nonlinear Risks: Mathematics Does Not Need to Rush

Original video: Terence Tao: SAIR’s Open Math Model Initiative

SAIR · 2026-09-18 · 25 minutes 15 seconds

Speaker: Terence Tao, Professor of Mathematics at UCLA, Fields Medalist, Co-founder of SAIR

Supporting primary sources: Declaration by 25 Fields Medalists · SAIR Open Math Model homepage

On 2026-09-11, the Science x AI Summit opened at Caltech, co-organized by SAIR, Caltech, and the Merkin Center for Pure and Applied Mathematics. Terence Tao delivered the opening keynote address. This article is based on the complete recorded address and verified subtitles.

On the lecture stage at Caltech, Terence Tao remarked that he had just experienced the strangest week of his life. That morning, together with 24 fellow Fields Medalist colleagues, he had published a joint declaration warning of a severe misalignment between frontier artificial intelligence and mathematics.

Commercial laboratories are generating answers faster than ever, but mathematics has encountered an unprecedented bottleneck. Artificial intelligence generates and verifies formal proofs rapidly, yet cannot replace the human work of digesting, explaining, and passing down mathematical understanding.

Beyond Problem Solving: The Complete Mathematical Lifecycle

Pure mathematics is basic, curiosity-driven research where major open conjectures serve as lighthouses. Navigators do not travel to lighthouses to live in them; rather, lighthouses illuminate the surrounding waterways and reefs so ships can navigate safely.

Solving an open problem involves a long and deliberate cognitive lifecycle. A proposed solution is only the first step. Early proofs are often messy and unverified, requiring careful checking by colleagues before being accepted as correct.

The next essential step is exposition. Raw, unrefined arguments must be rewritten clearly so other mathematicians can grasp the underlying principles and core insights.

Papers then enter peer review and formal publication, securing recognition from the wider community. This stage rarely attracts glamour; mathematics awards medals for solving problems rather than refereeing papers, yet thorough refereeing upholds the foundation of trust.

The final stage is canonicalization. Disparate papers are synthesized into coherent textbooks, enabling the next generation of students to comprehend connections across the broader discipline. Frontier language models perform competently in algebra and group theory precisely because centuries of human scholars curated these foundational textbooks.

"Proof Indigestion" and the Understanding Bottleneck

The current systemic imbalance arises because AI excels at generating candidate proofs and executing formal verification in languages such as Lean, while leaving exposition, publication, and pedagogical canonicalization virtually unchanged.

This divergence produces what Tao terms "proof indigestion." Immense volumes of machine-generated theorems pass Lean verification, yet even the engineers who initiated the runs cannot explain the underlying mechanisms or insights.

Tao illustrated the predicament through the evolution of urban transportation. In the early twentieth century, city streets built for pedestrians and horses were suddenly inundated with automobiles, producing citywide gridlock. The ideal future requires two parallel channels: dedicated freeways and transit corridors where AI performs high-throughput calculations at scale, alongside pedestrian zones and car-free parks where human researchers can think, converse, and discover at a deliberate pace.

Current industry practice resembles heavy steamrollers flattening the terrain. Frontier labs announce solutions to prominent conjectures, yet long sequences of formal symbols rarely translate into digestible knowledge. Pure mathematics faces no deadline next week, and compressing research cycles blindly risks eroding the discipline's core value.

25 Fields Medalists on AI Misalignment: When Grading Replaces Understanding

The declaration released on the day of the keynote at mathandai.org gathered signatures from 25 Fields Medalists, including Tao. The statement addresses commercial labs over-optimizing for easily graded benchmarks.

Frontier organizations concentrate on what automated systems can grade. A formal proof checker delivers an unambiguous pass or fail mark, whereas exposition, conceptual insight, and pedagogical clarity resist simple metrics.

When optimization pressure remains modest, measurable benchmarks align reasonably well with genuine scientific progress. Once extreme optimization takes over, gradable metrics decouple almost entirely from meaningful understanding.

The academic community holds considerable soft power. Technology companies pursue famous mathematical conjectures because they seek the prestige and rigor associated with mathematics. By actively articulating what mathematics values, researchers can guide frontier laboratories toward healthier technical priorities.

SAIR’s Open Math Model: Returning Tools to Researchers

A year ago, Tao co-founded the non-profit foundation SAIR alongside Chuck. In response to proprietary frontier systems, SAIR is organizing an open math model initiative built collaboratively with academic institutions and the open-source ecosystem.

The initiative aims to provide models whose weights researchers can inspect, run locally, adapt, and integrate into mathematics education. SAIR is assembling partners spanning compute infrastructure, research funding, and academic departments to cultivate responsible scientific computing.

Proprietary models package mathematical inquiry into closed digital black boxes. Researchers cannot audit internal reasoning traces, nor can they embed opaque services safely into long-term academic curricula.

An open math model provides working mathematicians with a transparent, verifiable instrument. Scholars can examine deduction steps rather than accepting black-box outputs, turning models into scaffolding for human intellect.

The Distillation Challenge: How a 10KB Prompt Boosted Open Models by 30% to 40%

Alongside long-term model development, SAIR established focused benchmark challenges to probe human-machine collaboration. The inaugural competition arose from Tao's Equational Theories project.

That initiative cataloged 22,000,000 algebraic true-or-false problems formalized in Lean, with each problem requiring roughly 30 minutes of graduate-level deduction. Frontier proprietary models answered 99% correctly after 5 minutes of internal deliberation, whereas local open-weight models hovered around 50%, equivalent to a random coin flip.

The initial stage challenged entrants to supply a prompt restricted to 10KB, functioning as a single-page reference cheat sheet for an algebra midterm. Well-engineered cheat sheets improved open-source performance by 30% to 40%, lifting overall success rates to 80% and 90%.

The newly completed Stage 2 expanded the submission budget to a 100KB Python script tasked with generating complete formal Lean proofs. Post-competition analysis yielded an unexpected finding: the highest-scoring teams made minimal calls to large language models, demonstrating that sound classical algorithm design frequently surpasses brute-force model queries.

The Inverse Galois Challenge: From Competition to a Collaborative Super-Team

A second major endeavor focused on the Inverse Galois problem. For degree 24 polynomials, theory predicts approximately 165,000 realizable Galois group signatures, of which mathematicians had previously identified only around 600, representing less than 0.3% coverage.

The first stage ran as an open Easter egg hunt, inviting worldwide participants to construct matching polynomials. Entrants identified 98.9% of all possible signatures, eliminating vast computational blind spots.

Notably, the team that outperformed all other contestants consisted of 2 human mathematicians. They relied on traditional computational algebra techniques, completing their work with virtually no reliance on artificial intelligence.

Stage 2 transitions from competition into open community collaboration. All participants share their methods in a single workspace, allowing an AI orchestration system to combine collective insights into a super-team targeting the remaining 1.1% of elusive signatures and advancing toward degree 31 polynomials.

Data Pollution and Nonlinear Risks: Mathematics Does Not Need to Rush

Addressing questions regarding models trained repeatedly on synthetic mathematical artifacts, Tao warned against the perils of data pollution. Much like biological cloning experiments where mice cloned across successive generations eventually lose viability, iterative AI loops risk cognitive degeneration.

Modern AI succeeds because it draws on centuries of curated, high-quality human literature. If unexamined synthetic outputs flood academic repositories and train future architectures, the signal-to-noise ratio will deteriorate.

Mathematical experience demonstrates that real-world systems are overwhelmingly nonlinear. Assuming that scaling compute by 10 or 100 will deliver proportional gains ignores how complex systems destabilize under runaway feedback.

Researchers must set their own cadence in an era of rapid automation. Rushing through problems without comprehension cannot yield lasting mathematical breakthroughs. Preserving lucid exposition, rigorous peer discourse, and conceptual depth remains the only durable path forward.

Published from
atlasnote-editorial
Published
2026-09-21
Tags
AImathematicsreasoningResearchopen-source