A formal proof checker receives a derivation and returns accept or reject. It has no input for whether the theorem was worth proving. Mathematics contains proofs, conjectures, explanations, examples, definitions, research programs, and the failed attempts by which students learn the subject. These objects have different feedback costs. Some can be checked almost immediately; others require taste, later use, experiments, or the judgment of a community. The question “is AI good at mathematics?” is therefore too coarse. A better unit is the task, together with the cost of evaluating a proposed answer.
AI systems move fastest when many candidate answers can be scored cheaply. A unit test can reject a code edit; a proof checker can reject a formal derivation; a benchmark can rank model outputs; a simulation can test a proposed structure or design. The loop generates candidates, rejects failures, keeps the better attempts, and searches again. Domain knowledge remains part of the search; failed attempts become cheaper. Current work on AI-assisted mathematics already uses this pattern: language-model search is coupled to proof checkers, formal systems, or other external tests (Avigad 2026; Tsoukalas et al. 2026). David Bessis describes theorem-proving as a public currency for mathematical value. When proof production becomes cheaper than concept-building, explanation, and canonization, that currency no longer measures the same bundle of work (Bessis 2026).
Proof Search
Formal proof search isolates one part of mathematical work. A proof has to be found, checked, and digested: finding it is search, checking it is certification, and digesting it turns the argument into something people can verify without heroic effort, teach to others, modify in nearby settings, and recognize in another problem. In the recent disproof of the unit distance conjecture, AI-assisted search produced mathematical material, and mathematicians then turned it into a short, human-verified account (Alon et al. 2026). In a digested proof, the construction has a name, the brittle case split has often been replaced by a lemma, and a reader can see which part of the argument might transfer. Older discussions of proof, understanding, and social verification already distinguish a correct proof from a usable piece of mathematics (Thurston 1994; De Millo, Lipton, and Perlis 1979). If AI increases the supply of correct arguments faster than mathematicians can digest them, the bottleneck moves from proof production to proof assimilation.
What A Proof Checker Cannot Check
Self-play and reinforcement learning need a reward. Games supply the reward through their rules (Silver et al. 2017). Formal correctness can play a related role. Research mathematics has too many correct statements. One can vary constants, add irrelevant hypotheses, prove narrow variants, or close lemmas that never compress a later argument. For example, a proof checker can certify a lemma after replacing a constant by a worse one and adding a boundedness assumption; the certificate does not say whether anyone should remember the variant. A theorem earns attention when it opens a method, explains an example, connects two areas, or makes later questions easier to ask. There is no cheap oracle for that judgment. A checker can certify that a derivation follows from the rules; it cannot certify that the theorem was worth proving. Taste enters as an expensive and delayed evaluation problem: the relevant labels arrive late and require expertise. As proof search improves, more value moves toward choosing statements, definitions, examples, and research programs that make the search worth running (Avigad 2026; Schwer 2026).
Synthesis Debt
A field has synthesis debt when it proves results faster than it turns them into usable understanding. The proof pays one cost; later users pay another when they have to teach the result, reuse it, compare it with older work, or reorganize a part of the literature around it. AI can increase the production side: more lemmas, more counterexamples, more formal derivations, more candidate proofs. The digestion side remains slower because it involves compression, exposition, comparison, teaching, and reuse. An undigested theorem can still close a question or block a false conjecture. A result can settle a problem while leaving no standard example, no named construction, and no short route through the proof. If many results accumulate before anyone has turned them into a navigable map, the field becomes harder to use even as the archive becomes larger. The unit-distance account reduced this debt by making the result easier to check, explain, and place in the existing literature (Alon et al. 2026; Avigad 2026; Bessis 2026).
Communities And Training
Documents record definitions, theorems, and arguments. Communities carry the working knowledge around them: which proof is standard, which example is misleading, which trick transfers, which statement is a dead end, and which reformulation is worth trying (Thurston 1994; De Millo, Lipton, and Perlis 1979). If routine mathematical work becomes automated, fewer people may pass through the apprenticeship that builds this knowledge. The papers may remain while fewer people know how to use them well. An expert and a novice extract different value from the same AI output. The expert has internal checks: they can reject a bad proof sketch, notice a missing hypothesis, or see that an argument proves the wrong thing. The novice may receive the same output while skipping the attempts and repairs that would have built those checks. For a student, asking for the proof before trying small cases can remove the moment where a missing hypothesis, a boundary case, or a bad definition first becomes visible. The educational problem is the loss of exercises that form judgment: checking, posing questions, comparing approaches, explaining failures, and deciding which results matter (Commelin et al. 2026; Schwer 2026; Klowden and Tao 2026).
When The Scoreboard Becomes The Target
Training follows incentives, and incentives follow what institutions can observe. Benchmarks and verifiers make progress visible (Abouzaid et al. 2026; Campbell 1979). Over time, the visible quantity can become the quantity that gets funded, hired for, and trained toward. AI produces gains where the score is clear, so the scored task can absorb attention. A benchmark can measure progress and still distort behavior once it becomes the easiest public measure of progress. If a public scoreboard counts formalized theorem statements solved, problems with clean statements, existing libraries, and fast checker feedback become especially attractive. Work whose hard part is finding the right definition or organizing a body of examples becomes harder to display on the same board. In mathematics and science, tractable work can start to look like high-value work because it is the part that moves on the scoreboard. When a community repeatedly rewards what is easy to score, young researchers acquire the taste of that reward system. Technical benchmarks therefore interact with the social organization of mathematics (Commelin et al. 2026).
The Conjecture Economy
Cheap proof raises the relative value of good questions. The hard work moves upstream: choosing the problem, finding the right definition, deciding which analogy is worth developing, and seeing why a statement should organize later work. Erdos-style conjectures have the relevant form: many are easy to state, hard to prove, and well chosen. If AI proves many such conjectures, the achievement of asking them becomes more visible (Tsoukalas et al. 2026). The proof closes the problem; the conjecture selected a problem worth closing. When machines can close many supplied statements, the scarce human act becomes selecting statements that were worth supplying. Proof remains valuable. Its relative price changes compared with question selection, synthesis, and explanation. Cheaper proofs make question selection, language-building, and explanation more visible (Avigad 2026; Alon et al. 2026).
A Feedback Cost Map
Feedback cost also applies outside mathematics. Field names are too coarse because feedback loops live inside fields. Biology contains fast-feedback structure prediction and slow-feedback clinical intervention; mathematics contains formal proof checking and open-ended conjecture formation; physics contains simulation-tight subproblems and theory choices that may wait years for evidence (Jumper et al. 2021; Tsoukalas et al. 2026). AlphaFold-style structure prediction can be checked against held-out structures; a proposed clinical intervention has to pass through patients, protocols, confounders, and time. Economics and climate contain many questions where controlled feedback is slow, noisy, or unavailable. Candidate answers become training data only when they can be evaluated reliably. Closed feedback loops let AI compound. Slow or unavailable feedback leaves progress to other routes. Examples from formal mathematics should therefore be generalized at the level of feedback loops. The department name is too coarse (Alon et al. 2026).
Plateau, AGI, And Labor
There is no single plateau question. Progress can continue where checks are cheap while slowing where feedback is absent, delayed, or expensive. A system’s abilities depend in part on the feedback loops it can close. More inference-time search helps directly when candidate answers can be scored. It helps less directly when the task is to choose a problem, invent a definition, or decide which model of a messy system is worth trusting. A system can therefore be strong at proof search and weak at choosing the problem. One system can occupy different positions on those tasks because the feedback costs differ.
Feedback cost also changes labor. If the human in an AI-assisted process is only a serial review bottleneck, there is economic pressure to reduce that bottleneck. Amdahl’s law imposes the corresponding constraint for knowledge work: speedup remains limited while every output waits for a slow human approval step. For example, an AI can draft many proof sketches, code patches, or reports in parallel; publication or deployment still waits when one expert has to inspect each item serially. Roles tied to question selection, responsibility, interpretation, and standards sit partly outside the cheap-feedback pipeline. Their durability comes from the objective-setting role itself, and no fixed boundary guarantees that the role remains human forever. They are harder to remove because removing them changes the objective being optimized. AI already does parts of mathematics and science. The remaining question is which feedback loops it can close, and what happens to the valuable work those loops do not measure (Avigad 2026; Klowden and Tao 2026; Gwern 2025).
Sources
Bessis describes the theorem economy: theorem-proving functions as a public currency for mathematical value, although concept-building, canonization, and intelligibility carry much of the value (Bessis 2026). Thurston and DeMillo-Lipton-Perlis describe proof as a social and explanatory object, beyond its role as a formal certificate (Thurston 1994; De Millo, Lipton, and Perlis 1979). Avigad, Klowden-Tao, Commelin-Jamnik-Ochigame-Taelman-Venkatesh, and Schwer discuss AI-assisted mathematics, mathematical understanding, education, and community practice (Avigad 2026, 2025; Klowden and Tao 2026; Commelin et al. 2026; Schwer 2026). First Proof and the AI-driven formal proof search paper provide current benchmark and formal-search examples (Abouzaid et al. 2026; Tsoukalas et al. 2026). The unit-distance note records AI-assisted search followed by human compression (Alon et al. 2026). Campbell states the background warning about measured proxies reshaping the activity being measured (Campbell 1979). Gwern supplies the Amdahl’s-law pressure from assistant workflows toward more autonomous systems (Gwern 2025).