Section Paragraph Plans

This file is the main drafting scaffold. Each section has a function, required content, source anchors, and an exit condition. The later blog draft should follow this structure unless the author changes the article spine.

Opening: The Wrong Question

Function: Replace the broad question “is AI good at math?” with the feedback-cost question.

Target length: 2 to 3 paragraphs.

Required content:

  • Mention that AI-for-math commentary often treats mathematics as one task.
  • State that the relevant split is between tasks with cheap checks and tasks without them.
  • Avoid saying the public discussion is stupid or shallow.

Source anchors:

  • No citation needed in the first paragraph unless a concrete source is mentioned.
  • Later cite Avigad or Alon et al. when the proof example begins.

Allowed opening move:

Start from the fact that a proof checker and a conjecture are different objects. One can be checked mechanically; the other requires a judgment about what is worth asking.

Exit condition: The reader understands that the article will classify AI progress by feedback cost.

Section 1: Cheap Feedback

Function: Define feedback cost.

Target length: 3 to 4 paragraphs.

Required content:

  • Define a cheap feedback loop: candidate, check, selection, update.
  • Give examples: proof checker, unit test, benchmark, simulation.
  • Explain why more sampling and more search help in that regime.
  • State that human difficulty is not the right predictor.

Source anchors:

  • Avigad for AI proving research-level theorems.
  • Alon et al. for the unit-distance case.

Avoid:

  • A long list of domains.
  • Saying “verifiable register” in the main prose unless first defined plainly.

Exit condition: The reader is ready to see proof as a clean case of cheap feedback.

Section 2: Proof Is Three Tasks

Function: Separate production of a proof from mathematical understanding.

Target length: 4 to 5 paragraphs.

Required content:

  • Define generation, verification, and digestion.
  • Use a short example of a long proof or AI proof needing human compression.
  • Use the unit-distance paper as the anchor.
  • State why digestion matters: teaching, reusing, recognizing special cases, moving techniques.

Source anchors:

  • Alon et al. for “short, digested, human-verified” version of the unit-distance argument.
  • Avigad for mathematical understanding and AI practice.

Avoid:

  • Saying machine proofs do not count.
  • Overusing “understanding” without specifying operations.

Exit condition: The reader accepts that correctness is not enough for cumulative mathematics.

Section 3: The Importance Oracle Is Missing

Function: Explain why proof search does not automatically produce good research.

Target length: 3 to 4 paragraphs.

Required content:

  • Self-play needs a reward.
  • Correctness can sometimes be the reward.
  • Importance is not cheaply scored.
  • Most true statements are not worth proving.
  • The human role shifts toward choosing statements, definitions, and programs.

Source anchors:

  • Avigad.
  • Schwer.
  • Internal source note for closed-objective-prerequisite-for-self-play.

Avoid:

  • Calling taste mysterious.
  • Claiming RLHF cannot help any part of research.

Exit condition: The reader sees why “more proofs” and “better mathematics” can diverge.

Section 4: Synthesis Debt

Function: State the production-versus-assimilation problem.

Target length: 3 to 4 paragraphs.

Required content:

  • Define synthesis debt.
  • Explain how AI increases production.
  • Explain why digestion remains costly.
  • State the failure mode: the archive grows faster than the map of the archive.

Source anchors:

  • Alon et al.
  • Avigad.
  • Internal source note for synthesis-debt-knowledge-overhang.

Avoid:

  • “Epistemic crisis.”
  • Treating all formal results as equal.

Exit condition: The reader is ready to move from individual proofs to communities.

Section 5: The Community Layer

Function: Explain why documents are not enough.

Target length: 3 to 4 paragraphs.

Required content:

  • State what documents contain.
  • State what communities provide: taste, warnings, standard examples, transfer knowledge.
  • Explain the risk if automation reduces the number of people trained in a field.
  • Keep this as a risk, not a prediction of collapse.

Source anchors:

  • Commelin et al.
  • Schwer.

Avoid:

  • Nostalgia.
  • “Civilization” language.

Exit condition: The reader is ready for the training problem.

Section 6: Acquisition Asymmetry

Function: Explain why the same tool has different effects on experts and students.

Target length: 4 to 5 paragraphs.

Required content:

  • Experts already have internal checks.
  • Novices may lose the exercise that builds those checks.
  • This is not mainly a cheating story.
  • The future issue is who will have judgment when AI is more capable.
  • End with the curriculum question: what should students practice if execution becomes cheap?

Source anchors:

  • Commelin et al.
  • Klowden and Tao.
  • Tao EMS webinar only after verification.

Avoid:

  • Blaming students.
  • Saying AI should be banned from education.

Exit condition: The reader is ready for institutional incentives.

Section 7: Measurability Distortion

Function: Show how tools reshape what institutions reward.

Target length: 4 paragraphs.

Required content:

  • Benchmarks start as proxies.
  • Proxies become targets.
  • AI produces visible gains where proxies are sharp.
  • Funding, hiring, and training can follow the visible gains.

Source anchors:

  • Commelin et al.
  • Leiden Declaration only after verification.

Avoid:

  • Saying all benchmarks are bad.
  • A generic rant about academia.

Exit condition: The reader is ready for the value shift from answers to questions.

Section 8: The Conjecture Economy

Function: Explain the incentive shift when proof gets cheaper.

Target length: 4 paragraphs.

Required content:

  • State the value shift: from proving supplied statements toward asking good questions.
  • Use the Erdos conjecture framing carefully.
  • Explain why prizes and promotions mostly see solved problems.
  • State the mismatch without proposing a full reform program.

Source anchors:

  • Alon et al.
  • Avigad.
  • MathOverflow only after verification.

Avoid:

  • Saying proof loses value.
  • Heroic language about question-makers.

Exit condition: The reader is ready to generalize from mathematics to science.

Section 9: The Feedback Cost Map Of Science

Function: Generalize the argument beyond mathematics.

Target length: 4 to 5 paragraphs.

Required content:

  • Explain that fields are mixed bags.
  • Give a compact map: formal systems and code; fast simulations; lab biology and clinical trials; macroeconomics and climate; theory choice and taste.
  • For each item, state the feedback cost, not just the field name.
  • Avoid precise timelines unless sourced.

Source anchors:

  • Internal synthesis from ../concepts/feedback_cost_map.md.
  • Max Welling only after verification.
  • Alon et al. as the formal-math example.

Avoid:

  • “Biology will be disrupted” without specifying the task.
  • Ranking entire disciplines.

Exit condition: The reader is ready for plateau and AGI to be reframed.

Section 10: Plateau, AGI, And Labor

Function: Close the article by applying the feedback-cost split to broad AI questions.

Target length: 5 to 6 paragraphs.

Required content:

  • Split the plateau question into engineering limits and feedback limits.
  • State that progress can continue in cheap-check domains while remaining slower elsewhere.
  • Explain why AGI is too coarse for this essay’s question.
  • Bring in labor: assistant workflows still have a human bottleneck.
  • State what is harder to remove: objective setting, responsibility, taste, interpretation.
  • End soberly by returning to feedback loops.

Source anchors:

  • Avigad.
  • Klowden and Tao.
  • Gwern for Amdahl’s-law and assistant-to-replacement pressure.

Avoid:

  • Any fixed timeline.
  • “AGI is meaningless” as a blanket claim.
  • A dark ending.

Exit condition: The final paragraph restates the article’s practical test: which feedback loop can AI close?