Article Blueprint

This is the paragraph-level plan for the post. It should be compact enough to draft from directly. The target is one substantial paragraph per body section.

Drafting Assumption

Draft 1 should use only verified sources from source_map.md. Unverified references can remain in scaffold notes, but should not appear as public evidence in the post.

Working Title

AI Scales Where Feedback Is Cheap

This is a working title, not a final decision.

Target Shape

  • Opening: 1 to 2 paragraphs.
  • Body: 8 to 10 sections.
  • Each body section: 1 substantial paragraph by default.
  • Split a body section only when the example would otherwise crowd out the mechanism.
  • Total target: about 11 to 14 paragraphs.

Opening: The Wrong Question

Paragraphs: 1 to 2

Job: Replace “is AI good at mathematics?” with “which parts of mathematical work have cheap feedback?”

Core claim: Mathematics is not one task. A proof checker, a conjecture, a proof explanation, and a research program have different feedback structures.

Concrete example: A proof checker can verify a formal derivation; it cannot decide that a conjecture is worth attention.

Source: No citation needed in the opening unless a source is named.

Handoff: Introduce feedback cost as the article’s organizing variable.

Do not say: Online commentary is shallow; everyone is asking the wrong question.

Cheap Feedback

Paragraphs: 1

Job: Define the mechanism.

Core claim: AI systems improve fastest when candidate outputs can be checked cheaply and repeatedly. A cheap feedback loop has four parts: generate candidates, check them, keep the successful ones, and use the result to search better.

Concrete examples: Unit tests for code; proof checkers for formal proofs; benchmarks for fixed tasks; simulations for some scientific models.

Source: Avigad for AI-assisted mathematical practice. Alon et al. later for the unit-distance example.

Handoff: Proof search is the clean mathematical case because correctness can sometimes be checked.

Do not say: Cheap feedback is the only variable, or the verifier solves the whole problem.

Proof Is Three Tasks

Paragraphs: 1

Job: Separate proof production from mathematical understanding.

Core claim: Proof has at least three tasks: generation, verification, and digestion. AI may help strongly with the first two while leaving much of the third to humans.

Concrete example: The unit-distance case: AI-assisted search produced material that was later turned into a short, human-verified mathematical account.

Source: Alon et al. for the unit-distance case. Avigad for mathematical understanding and AI practice.

Handoff: Correctness is not enough for cumulative mathematics; the next issue is which statements are worth proving.

Do not say: Machine proofs do not count, or only digestion is real mathematics.

The Importance Oracle

Paragraphs: 1

Job: Explain why proof search does not automatically produce good research.

Core claim: Self-play needs a reward. Correctness can sometimes be checked, but importance cannot be scored cheaply. Most true statements are not worth proving.

Concrete example: In a game, the rules supply win/loss feedback. In mathematics, a true lemma may be irrelevant unless it compresses, connects, or opens further work.

Source: Avigad. Schwer. Internal source note for closed-objective-prerequisite-for-self-play.

Handoff: If more true statements can be produced, the field needs more synthesis.

Do not say: AI will never have taste, or RLHF is useless.

Synthesis Debt

Paragraphs: 1

Job: State the production-versus-assimilation problem.

Core claim: A field has synthesis debt when it proves results faster than it turns them into usable understanding. AI can increase theorem production faster than the community can compress, explain, and reuse the results.

Concrete example: Long machine-generated proofs or technical artifacts that close questions but require human cleanup before they become useful.

Source: Alon et al. Avigad. Internal source note for synthesis-debt-knowledge-overhang.

Handoff: This makes the community around a field more important, not less.

Do not say: Theorem counts are meaningless, or all AI-generated proofs are opaque.

The Community Layer

Paragraphs: 1

Job: Explain why documents are not enough.

Core claim: Documents record arguments, but communities carry much of the knowledge needed to use them. A community knows which proofs matter, which tricks transfer, which examples are misleading, and which questions are exhausted.

Concrete example: An area can have complete papers but few active experts, making it difficult for outsiders to enter or reuse.

Source: Commelin et al. Schwer.

Handoff: The next issue is training: who will still acquire this working knowledge?

Do not say: Documents are unimportant, or AI will kill mathematical communities.

Acquisition Asymmetry

Paragraphs: 1

Job: Explain why the same tool has different effects on experts and students.

Core claim: An expert can use AI while relying on already-formed judgment. A novice may bypass the failed attempts, repairs, and comparisons that build that judgment. The issue is training, not cheating.

Concrete example: An expert can reject a bad proof sketch; a student may never build the skill needed to see why it is bad.

Source: Commelin et al. Klowden and Tao.

Handoff: Training follows incentives; institutions shape what gets practiced and rewarded.

Do not say: Students should not use AI, or correct answers are educationally worthless.

Measurability Distortion

Paragraphs: 1

Job: Show how tools reshape what institutions reward.

Core claim: Institutions tend to reward work that is easy to observe, compare, and count. AI strengthens this loop because it produces visible gains where the proxy is sharp.

Concrete example: Benchmarks begin as proxies, become targets, and then influence funding, hiring, and training.

Source: Commelin et al. Leiden Declaration only after verification.

Handoff: If proof and benchmarked progress become easier to obtain, good questions become relatively more valuable.

Do not say: Benchmarks are bad, institutions are stupid, or all measured work is shallow.

The Conjecture Economy

Paragraphs: 1

Job: Explain the value shift when proof gets cheaper.

Core claim: If proof becomes easier to obtain, good questions become relatively more valuable. The hard part moves upstream: choosing the problem, formulating it well, and seeing why it should organize future work.

Concrete example: If AI proves many Erdos-style conjectures, that emphasizes the value of the question-making. The proof closes the problem; the conjecture selected a problem worth closing.

Source: Avigad. Alon et al. MathOverflow only after verification.

Handoff: This feedback-cost view also applies outside mathematics.

Do not say: Proof stops mattering, or the prize system is simply wrong.

Feedback Cost Map Of Science

Paragraphs: 1

Job: Generalize beyond mathematics.

Core claim: Disciplines are too coarse. AI disruption should be predicted by feedback cost: formal systems and code differ from clinical trials, macroeconomics, climate policy, and theory choice.

Concrete example: Biology contains fast-feedback structure prediction and slow-feedback clinical intervention. Mathematics contains formal proof checking and open-ended conjecture formation.

Source: Alon et al. for the formal-math anchor. Max Welling only after verification. Internal synthesis from feedback_cost_map.md.

Handoff: This reframes plateau and AGI.

Do not say: Whole fields are safe or doomed, or precise timelines without support.

Plateau, AGI, And Labor

Paragraphs: 1

Job: Close by applying the feedback-cost split to broad AI questions.

Core claim: There is no single plateau question. Progress can continue where feedback is cheap while remaining slower where feedback is absent, delayed, or expensive. For this article, the useful AGI question is which feedback loops a system can close.

Concrete example: More inference-time search helps proof or code more directly than it helps theory choice with no clear reward. A system can be strong at proof search and weak at choosing the problem.

Labor claim: If the human remains only a serial review bottleneck, the economic pressure is to reduce that bottleneck. Roles tied to question selection, responsibility, interpretation, and standards are harder to remove than labor inside a cheap-feedback pipeline.

Source: Avigad. Klowden and Tao. Gwern for the assistant-to-replacement pressure.

Ending: Return to the feedback-loop test: AI already does parts of mathematics and science. The useful question is which feedback loops it can close, and what happens to the work those loops do not measure.

Do not say: There is definitely no plateau; AGI is meaningless in every context; all researchers will be replaced; grand claims about the fate of society.