Section Paragraph Plans
This file is the main drafting scaffold. Each section has a function, required content, source anchors, and an exit condition. The later blog draft should follow this structure unless the author changes the article spine.
Opening: The Wrong Question
Function: Replace the broad question “is AI good at math?” with the feedback-cost question.
Target length: 2 to 3 paragraphs.
Required content:
- Mention that AI-for-math commentary often treats mathematics as one task.
- State that the relevant split is between tasks with cheap checks and tasks without them.
- Avoid saying the public discussion is stupid or shallow.
Source anchors:
- No citation needed in the first paragraph unless a concrete source is mentioned.
- Later cite Avigad or Alon et al. when the proof example begins.
Allowed opening move:
Start from the fact that a proof checker and a conjecture are different objects. One can be checked mechanically; the other requires a judgment about what is worth asking.
Exit condition: The reader understands that the article will classify AI progress by feedback cost.
Section 1: Cheap Feedback
Function: Define feedback cost.
Target length: 3 to 4 paragraphs.
Required content:
- Define a cheap feedback loop: candidate, check, selection, update.
- Give examples: proof checker, unit test, benchmark, simulation.
- Explain why more sampling and more search help in that regime.
- State that human difficulty is not the right predictor.
Source anchors:
- Avigad for AI proving research-level theorems.
- Alon et al. for the unit-distance case.
Avoid:
- A long list of domains.
- Saying “verifiable register” in the main prose unless first defined plainly.
Exit condition: The reader is ready to see proof as a clean case of cheap feedback.
Section 2: Proof Is Three Tasks
Function: Separate production of a proof from mathematical understanding.
Target length: 4 to 5 paragraphs.
Required content:
- Define generation, verification, and digestion.
- Use a short example of a long proof or AI proof needing human compression.
- Use the unit-distance paper as the anchor.
- State why digestion matters: teaching, reusing, recognizing special cases, moving techniques.
Source anchors:
- Alon et al. for “short, digested, human-verified” version of the unit-distance argument.
- Avigad for mathematical understanding and AI practice.
Avoid:
- Saying machine proofs do not count.
- Overusing “understanding” without specifying operations.
Exit condition: The reader accepts that correctness is not enough for cumulative mathematics.
Section 3: The Importance Oracle Is Missing
Function: Explain why proof search does not automatically produce good research.
Target length: 3 to 4 paragraphs.
Required content:
- Self-play needs a reward.
- Correctness can sometimes be the reward.
- Importance is not cheaply scored.
- Most true statements are not worth proving.
- The human role shifts toward choosing statements, definitions, and programs.
Source anchors:
- Avigad.
- Schwer.
- Internal source note for
closed-objective-prerequisite-for-self-play.
Avoid:
- Calling taste mysterious.
- Claiming RLHF cannot help any part of research.
Exit condition: The reader sees why “more proofs” and “better mathematics” can diverge.
Section 4: Synthesis Debt
Function: State the production-versus-assimilation problem.
Target length: 3 to 4 paragraphs.
Required content:
- Define synthesis debt.
- Explain how AI increases production.
- Explain why digestion remains costly.
- State the failure mode: the archive grows faster than the map of the archive.
Source anchors:
- Alon et al.
- Avigad.
- Internal source note for
synthesis-debt-knowledge-overhang.
Avoid:
- “Epistemic crisis.”
- Treating all formal results as equal.
Exit condition: The reader is ready to move from individual proofs to communities.
Section 5: The Community Layer
Function: Explain why documents are not enough.
Target length: 3 to 4 paragraphs.
Required content:
- State what documents contain.
- State what communities provide: taste, warnings, standard examples, transfer knowledge.
- Explain the risk if automation reduces the number of people trained in a field.
- Keep this as a risk, not a prediction of collapse.
Source anchors:
- Commelin et al.
- Schwer.
Avoid:
- Nostalgia.
- “Civilization” language.
Exit condition: The reader is ready for the training problem.
Section 6: Acquisition Asymmetry
Function: Explain why the same tool has different effects on experts and students.
Target length: 4 to 5 paragraphs.
Required content:
- Experts already have internal checks.
- Novices may lose the exercise that builds those checks.
- This is not mainly a cheating story.
- The future issue is who will have judgment when AI is more capable.
- End with the curriculum question: what should students practice if execution becomes cheap?
Source anchors:
- Commelin et al.
- Klowden and Tao.
- Tao EMS webinar only after verification.
Avoid:
- Blaming students.
- Saying AI should be banned from education.
Exit condition: The reader is ready for institutional incentives.
Section 7: Measurability Distortion
Function: Show how tools reshape what institutions reward.
Target length: 4 paragraphs.
Required content:
- Benchmarks start as proxies.
- Proxies become targets.
- AI produces visible gains where proxies are sharp.
- Funding, hiring, and training can follow the visible gains.
Source anchors:
- Commelin et al.
- Leiden Declaration only after verification.
Avoid:
- Saying all benchmarks are bad.
- A generic rant about academia.
Exit condition: The reader is ready for the value shift from answers to questions.
Section 8: The Conjecture Economy
Function: Explain the incentive shift when proof gets cheaper.
Target length: 4 paragraphs.
Required content:
- State the value shift: from proving supplied statements toward asking good questions.
- Use the Erdos conjecture framing carefully.
- Explain why prizes and promotions mostly see solved problems.
- State the mismatch without proposing a full reform program.
Source anchors:
- Alon et al.
- Avigad.
- MathOverflow only after verification.
Avoid:
- Saying proof loses value.
- Heroic language about question-makers.
Exit condition: The reader is ready to generalize from mathematics to science.
Section 9: The Feedback Cost Map Of Science
Function: Generalize the argument beyond mathematics.
Target length: 4 to 5 paragraphs.
Required content:
- Explain that fields are mixed bags.
- Give a compact map: formal systems and code; fast simulations; lab biology and clinical trials; macroeconomics and climate; theory choice and taste.
- For each item, state the feedback cost, not just the field name.
- Avoid precise timelines unless sourced.
Source anchors:
- Internal synthesis from
../concepts/feedback_cost_map.md. - Max Welling only after verification.
- Alon et al. as the formal-math example.
Avoid:
- “Biology will be disrupted” without specifying the task.
- Ranking entire disciplines.
Exit condition: The reader is ready for plateau and AGI to be reframed.
Section 10: Plateau, AGI, And Labor
Function: Close the article by applying the feedback-cost split to broad AI questions.
Target length: 5 to 6 paragraphs.
Required content:
- Split the plateau question into engineering limits and feedback limits.
- State that progress can continue in cheap-check domains while remaining slower elsewhere.
- Explain why AGI is too coarse for this essay’s question.
- Bring in labor: assistant workflows still have a human bottleneck.
- State what is harder to remove: objective setting, responsibility, taste, interpretation.
- End soberly by returning to feedback loops.
Source anchors:
- Avigad.
- Klowden and Tao.
- Gwern for Amdahl’s-law and assistant-to-replacement pressure.
Avoid:
- Any fixed timeline.
- “AGI is meaningless” as a blanket claim.
- A dark ending.
Exit condition: The final paragraph restates the article’s practical test: which feedback loop can AI close?