Article Spine
This is the intended structure of the blog post. It is a scaffold, not prose.
Global Argument
AI progress in mathematics and science follows the cost of feedback. When candidate outputs can be checked cheaply, progress can compound. When the hard part is choosing the right question, digesting a result, training judgment, or maintaining a community that knows how to use the literature, the bottleneck moves elsewhere.
Section Order
Opening: The Wrong Question
Purpose: Replace “is AI good at math?” with “which parts of mathematical work have cheap feedback?”
Content:
- Start from AI-for-math commentary being too coarse.
- State that “math” is not one task.
- Introduce feedback cost as the organizing variable.
Exit: The reader is ready for the proof example.
Section 1: Cheap Feedback
Purpose: Define the mechanism.
Content:
- Candidate generation plus cheap evaluation gives a learning/search loop.
- Examples: proof checker, unit test, benchmark, simulation.
- In these settings, more sampling and better search can pay off.
Exit: The reader understands why proof search is a natural early case.
Section 2: Proof Is Three Tasks
Purpose: Separate generation, verification, and digestion.
Content:
- Generation: finding an argument.
- Verification: checking correctness.
- Digestion: compressing, explaining, teaching, connecting, reusing.
- Use the unit-distance example as the anchor.
Exit: The reader can see why a correct proof is not the same as usable mathematical knowledge.
Section 3: The Importance Oracle Is Missing
Purpose: Explain why self-play does not solve research taste.
Content:
- Self-play needs a reward.
- Correctness can sometimes be checked.
- Importance cannot be checked cheaply.
- Most true statements are not worth attention.
Exit: The reader understands why “prove more theorems” and “do better mathematics” can diverge.
Section 4: Synthesis Debt
Purpose: Explain the risk of producing results faster than they are assimilated.
Content:
- Define synthesis debt.
- AI can increase theorem production faster than digestion.
- Long or opaque proofs can close questions without creating much reusable understanding.
Exit: The reader is ready to think about the community practices around the documents.
Section 5: The Community Layer
Purpose: Show that mathematical knowledge depends on living practice.
Content:
- Documents are not enough.
- A community knows which proofs matter, which tricks transfer, which examples are misleading.
- If fewer people deeply learn an area, the archive can grow while operational knowledge weakens.
Exit: The reader is ready for the education problem.
Section 6: Acquisition Asymmetry
Purpose: Explain how AI can help experts while weakening novices.
Content:
- Expert use depends on already-formed judgment.
- Novice use can bypass the exercises that form judgment.
- The issue is not cheating; it is training.
- The AI-native generation may later need to direct more capable systems.
Exit: The reader is ready for institutional consequences.
Section 7: Measurability Distortion
Purpose: Explain how tools reshape what gets rewarded.
Content:
- Benchmarks and verifiers make some work easier to compare.
- Institutions tend to reward visible progress.
- AI can make tractable work look like the frontier.
- Training then adapts to the reward system.
Exit: The reader is ready for the conjecture economy.
Section 8: The Conjecture Economy
Purpose: Describe the value shift from closing problems to choosing them.
Content:
- If proof becomes cheaper, good questions become relatively more valuable.
- Erdos-style conjectures illustrate the point.
- Prizes and promotions mostly reward solved problems.
- The mismatch grows when solving becomes easier to automate.
Exit: The reader is ready to generalize beyond mathematics.
Section 9: The Feedback Cost Map Of Science
Purpose: Generalize from mathematics to science.
Content:
- Fields are too coarse.
- Feedback regimes cut across fields.
- Formal systems, code, and fast simulations differ from clinical trials, macroeconomics, and theory choice.
- Prediction should be by feedback cost, not by discipline name.
Exit: The reader is ready to revisit plateau and AGI.
Section 10: Plateau, AGI, And Labor
Purpose: Close by replacing broad AI questions with the feedback-cost partition.
Content:
- No single plateau question.
- Progress can continue where checks are cheap while remaining slower elsewhere.
- AGI is too coarse for this issue.
- Economically, assistant workflows may be unstable if the human is only a review bottleneck.
- Roles tied to question selection, responsibility, and interpretation are harder to remove.
Exit: The post ends with the main claim restated in practical terms.
Possible Ending
The ending should be sober, not dramatic. It should return to the feedback-cost split:
The question is not whether AI will do mathematics or science. It already does parts of both. The better question is which feedback loops it can close, and what happens to the parts of the discipline whose value was never captured by those loops.