AI and the Cost of Feedback

AI
mathematics
science
Published

03 07 2026

Modified

04 07 2026

A formal proof checker receives a derivation and returns accept or reject. It has no input for whether the theorem was worth proving. Mathematics contains proofs, conjectures, explanations, examples, definitions, research programs, and the failed attempts by which students learn the subject. These objects have different feedback costs. Some can be checked almost immediately; others require taste, later use, experiments, or the judgment of a community. The question “is AI good at mathematics?” is therefore too coarse. A better unit is the task, together with the cost of evaluating a proposed answer.

AI systems move fastest when many candidate answers can be scored cheaply. A unit test can reject a code edit; a proof checker can reject a formal derivation; a benchmark can rank model outputs; a simulation can test a proposed structure or design. The loop generates candidates, rejects failures, keeps the better attempts, and searches again. Domain knowledge remains part of the search; failed attempts become cheaper. Current work on AI-assisted mathematics already uses this pattern: language-model search is coupled to proof checkers, formal systems, or other external tests (Avigad 2026; Tsoukalas et al. 2026). David Bessis describes theorem-proving as a public currency for mathematical value. When proof production becomes cheaper than concept-building, explanation, and canonization, that currency no longer measures the same bundle of work (Bessis 2026).

What A Proof Checker Cannot Check

Self-play and reinforcement learning need a reward. Games supply the reward through their rules (Silver et al. 2017). Formal correctness can play a related role. Research mathematics has too many correct statements. One can vary constants, add irrelevant hypotheses, prove narrow variants, or close lemmas that never compress a later argument. For example, a proof checker can certify a lemma after replacing a constant by a worse one and adding a boundedness assumption; the certificate does not say whether anyone should remember the variant. A theorem earns attention when it opens a method, explains an example, connects two areas, or makes later questions easier to ask. There is no cheap oracle for that judgment. A checker can certify that a derivation follows from the rules; it cannot certify that the theorem was worth proving. Taste enters as an expensive and delayed evaluation problem: the relevant labels arrive late and require expertise. As proof search improves, more value moves toward choosing statements, definitions, examples, and research programs that make the search worth running (Avigad 2026; Schwer 2026).

Synthesis Debt

A field has synthesis debt when it proves results faster than it turns them into usable understanding. The proof pays one cost; later users pay another when they have to teach the result, reuse it, compare it with older work, or reorganize a part of the literature around it. AI can increase the production side: more lemmas, more counterexamples, more formal derivations, more candidate proofs. The digestion side remains slower because it involves compression, exposition, comparison, teaching, and reuse. An undigested theorem can still close a question or block a false conjecture. A result can settle a problem while leaving no standard example, no named construction, and no short route through the proof. If many results accumulate before anyone has turned them into a navigable map, the field becomes harder to use even as the archive becomes larger. The unit-distance account reduced this debt by making the result easier to check, explain, and place in the existing literature (Alon et al. 2026; Avigad 2026; Bessis 2026).

Communities And Training

Documents record definitions, theorems, and arguments. Communities carry the working knowledge around them: which proof is standard, which example is misleading, which trick transfers, which statement is a dead end, and which reformulation is worth trying (Thurston 1994; De Millo, Lipton, and Perlis 1979). If routine mathematical work becomes automated, fewer people may pass through the apprenticeship that builds this knowledge. The papers may remain while fewer people know how to use them well. An expert and a novice extract different value from the same AI output. The expert has internal checks: they can reject a bad proof sketch, notice a missing hypothesis, or see that an argument proves the wrong thing. The novice may receive the same output while skipping the attempts and repairs that would have built those checks. For a student, asking for the proof before trying small cases can remove the moment where a missing hypothesis, a boundary case, or a bad definition first becomes visible. The educational problem is the loss of exercises that form judgment: checking, posing questions, comparing approaches, explaining failures, and deciding which results matter (Commelin et al. 2026; Schwer 2026; Klowden and Tao 2026).

When The Scoreboard Becomes The Target

Training follows incentives, and incentives follow what institutions can observe. Benchmarks and verifiers make progress visible (Abouzaid et al. 2026; Campbell 1979). Over time, the visible quantity can become the quantity that gets funded, hired for, and trained toward. AI produces gains where the score is clear, so the scored task can absorb attention. A benchmark can measure progress and still distort behavior once it becomes the easiest public measure of progress. If a public scoreboard counts formalized theorem statements solved, problems with clean statements, existing libraries, and fast checker feedback become especially attractive. Work whose hard part is finding the right definition or organizing a body of examples becomes harder to display on the same board. In mathematics and science, tractable work can start to look like high-value work because it is the part that moves on the scoreboard. When a community repeatedly rewards what is easy to score, young researchers acquire the taste of that reward system. Technical benchmarks therefore interact with the social organization of mathematics (Commelin et al. 2026).

The Conjecture Economy

Cheap proof raises the relative value of good questions. The hard work moves upstream: choosing the problem, finding the right definition, deciding which analogy is worth developing, and seeing why a statement should organize later work. Erdos-style conjectures have the relevant form: many are easy to state, hard to prove, and well chosen. If AI proves many such conjectures, the achievement of asking them becomes more visible (Tsoukalas et al. 2026). The proof closes the problem; the conjecture selected a problem worth closing. When machines can close many supplied statements, the scarce human act becomes selecting statements that were worth supplying. Proof remains valuable. Its relative price changes compared with question selection, synthesis, and explanation. Cheaper proofs make question selection, language-building, and explanation more visible (Avigad 2026; Alon et al. 2026).

A Feedback Cost Map

Feedback cost also applies outside mathematics. Field names are too coarse because feedback loops live inside fields. Biology contains fast-feedback structure prediction and slow-feedback clinical intervention; mathematics contains formal proof checking and open-ended conjecture formation; physics contains simulation-tight subproblems and theory choices that may wait years for evidence (Jumper et al. 2021; Tsoukalas et al. 2026). AlphaFold-style structure prediction can be checked against held-out structures; a proposed clinical intervention has to pass through patients, protocols, confounders, and time. Economics and climate contain many questions where controlled feedback is slow, noisy, or unavailable. Candidate answers become training data only when they can be evaluated reliably. Closed feedback loops let AI compound. Slow or unavailable feedback leaves progress to other routes. Examples from formal mathematics should therefore be generalized at the level of feedback loops. The department name is too coarse (Alon et al. 2026).

Plateau, AGI, And Labor

There is no single plateau question. Progress can continue where checks are cheap while slowing where feedback is absent, delayed, or expensive. A system’s abilities depend in part on the feedback loops it can close. More inference-time search helps directly when candidate answers can be scored. It helps less directly when the task is to choose a problem, invent a definition, or decide which model of a messy system is worth trusting. A system can therefore be strong at proof search and weak at choosing the problem. One system can occupy different positions on those tasks because the feedback costs differ.

Feedback cost also changes labor. If the human in an AI-assisted process is only a serial review bottleneck, there is economic pressure to reduce that bottleneck. Amdahl’s law imposes the corresponding constraint for knowledge work: speedup remains limited while every output waits for a slow human approval step. For example, an AI can draft many proof sketches, code patches, or reports in parallel; publication or deployment still waits when one expert has to inspect each item serially. Roles tied to question selection, responsibility, interpretation, and standards sit partly outside the cheap-feedback pipeline. Their durability comes from the objective-setting role itself, and no fixed boundary guarantees that the role remains human forever. They are harder to remove because removing them changes the objective being optimized. AI already does parts of mathematics and science. The remaining question is which feedback loops it can close, and what happens to the valuable work those loops do not measure (Avigad 2026; Klowden and Tao 2026; Gwern 2025).

Sources

Bessis describes the theorem economy: theorem-proving functions as a public currency for mathematical value, although concept-building, canonization, and intelligibility carry much of the value (Bessis 2026). Thurston and DeMillo-Lipton-Perlis describe proof as a social and explanatory object, beyond its role as a formal certificate (Thurston 1994; De Millo, Lipton, and Perlis 1979). Avigad, Klowden-Tao, Commelin-Jamnik-Ochigame-Taelman-Venkatesh, and Schwer discuss AI-assisted mathematics, mathematical understanding, education, and community practice (Avigad 2026, 2025; Klowden and Tao 2026; Commelin et al. 2026; Schwer 2026). First Proof and the AI-driven formal proof search paper provide current benchmark and formal-search examples (Abouzaid et al. 2026; Tsoukalas et al. 2026). The unit-distance note records AI-assisted search followed by human compression (Alon et al. 2026). Campbell states the background warning about measured proxies reshaping the activity being measured (Campbell 1979). Gwern supplies the Amdahl’s-law pressure from assistant workflows toward more autonomous systems (Gwern 2025).

References

Abouzaid, Mohammed, Andrew J. Blumberg, Martin Hairer, Joe Kileel, Tamara G. Kolda, Paul D. Nelson, Daniel Spielman, et al. 2026. “First Proof.” arXiv preprint arXiv:2602.05192.
Alon, Noga, Thomas F. Bloom, W. T. Gowers, Daniel Litt, Will Sawin, Arul Shankar, Jacob Tsimerman, Victor Wang, and Melanie Matchett Wood. 2026. “Remarks on the Disproof of the Unit Distance Conjecture.” arXiv preprint arXiv:2605.20695.
Avigad, Jeremy. 2025. “Is Mathematics Obsolete?” arXiv preprint arXiv:2502.14874.
———. 2026. “Mathematicians in the Age of AI.” arXiv preprint arXiv:2603.03684.
Bessis, David. 2026. “The Fall of the Theorem Economy.” https://davidbessis.substack.com/p/the-fall-of-the-theorem-economy.
Campbell, Donald T. 1979. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning 2 (1). Elsevier BV: 67–90. doi:10.1016/0149-7189(79)90048-x.
Commelin, Johan, Mateja Jamnik, Rodrigo Ochigame, Lenny Taelman, and Akshay Venkatesh. 2026. “Shaping the Future of Mathematics in the Age of AI.” In Notices of the American Mathematical Society.
De Millo, Richard A., Richard J. Lipton, and Alan J. Perlis. 1979. “Social Processes and Proofs of Theorems and Programs.” Communications of the ACM 22 (5). Association for Computing Machinery (ACM): 271–80. doi:10.1145/359104.359106.
Gwern. 2025. “Guardian Angels: LLM Personalization for Productivity and Security.” https://gwern.net/guardian-angel.
Jumper, John, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, et al. 2021. “Highly Accurate Protein Structure Prediction with AlphaFold.” Nature 596 (7873). Springer Science; Business Media LLC: 583–89. doi:10.1038/s41586-021-03819-2.
Klowden, Tanya, and Terence Tao. 2026. “Mathematical Methods and Human Thought in the Age of AI.” arXiv preprint arXiv:2603.26524.
Schwer, Petra. 2026. “The Meaning of Doing Mathematics.” Mitteilungen Der Deutschen Mathematiker-Vereinigung 34 (1). Walter de Gruyter GmbH: 14–16. doi:10.1515/dmvm-2026-0008.
Silver, David, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, et al. 2017. “Mastering the Game of Go Without Human Knowledge.” Nature 550 (7676). Springer Science; Business Media LLC: 354–59. doi:10.1038/nature24270.
Thurston, William P. 1994. “On Proof and Progress in Mathematics.” Bulletin of the American Mathematical Society 30 (2). American Mathematical Society (AMS): 161–78. doi:10.1090/s0273-0979-1994-00502-6.
Tsoukalas, George, Anton Kovsharov, Sergey Shirobokov, Anja Surina, Moritz Firsching, Gergely Bérczi, Francisco J. R. Ruiz, et al. 2026. “Advancing Mathematics Research with AI-Driven Formal Proof Search.” arXiv preprint arXiv:2605.22763.