Article Flow Test

This is a prose test of the article’s movement. It is not the blog post. It checks whether the scaffold forms one argument rather than a sequence of adjacent ideas.

The article opens by changing the scale of the question. “Is AI good at mathematics?” treats mathematics as one activity, when the actual work contains many objects with different feedback structures: a formal derivation, a conjecture, a proof explanation, a research program, a student’s attempted solution. A proof checker can accept or reject a formal derivation, but it cannot tell us that a conjecture is worth ten years of attention. That contrast gives the opening its object. The useful variable is feedback cost. Some mathematical tasks supply cheap checks; others require taste, experiments, long-term use, or a community’s judgment before anyone knows whether the answer was good.

The first section defines the mechanism without turning it into a slogan. AI systems become especially useful when they can generate many candidates and receive cheap, reliable signals about which candidates worked. Unit tests do this for code; proof checkers do it for formal derivations; benchmarks do it for fixed tasks; simulations do it for some scientific models. In those settings, more sampling, better search, and more inference-time computation can produce visible progress because failures are cheap to discard. The section must keep the claim narrow. Cheap feedback does not make a task easy, and it does not solve the whole problem. It changes the economics of search.

Proof search then gives the clean mathematical case. The article separates three tasks that are often collapsed: generating a proof, verifying a proof, and digesting a proof. Generation finds an argument; verification checks whether the argument is correct; digestion turns the argument into something people can teach, reuse, modify, and recognize elsewhere. The unit-distance episode anchors this distinction because AI-assisted search produced material that later had to be converted into a short, human-verified mathematical account. The point is not to deny the proof or diminish the machine contribution. The point is that a correct proof and a usable piece of mathematics are different objects. If AI increases the supply of correct arguments faster than the community can digest them, the bottleneck has moved.

The importance problem follows from the same split. Self-play and reinforcement learning require a reward. In games, the reward is built into the rules; in many formal tasks, correctness can play that role. Research mathematics has a weaker signal. Most true statements are not worth much attention. A useful theorem compresses many cases, opens a method, connects two areas, explains an example, or makes later questions easier to ask. There is no cheap oracle for that judgment. The article should treat taste as an expensive and delayed evaluation problem, not as something mystical. As proof search improves, more value moves toward choosing statements, definitions, examples, and programs that make the search worth running.

The next turn is synthesis debt. A field can prove results faster than it assimilates them, and AI can widen that gap by increasing the production side: more lemmas, more counterexamples, more formal derivations, more candidate proofs. The digestion side remains slower because it involves compression, exposition, comparison, teaching, and reuse. An undigested theorem can still be real and useful; it can close a question or block a false conjecture. But if many results accumulate before anyone has turned them into a navigable map, the field becomes harder to use even as the archive becomes larger. The archive grows while working understanding grows more slowly.

From there the article moves from documents to communities. Documents record definitions, theorems, and arguments, but communities carry much of the working knowledge around them: which proof is standard, which example is misleading, which trick transfers, which statement is a dead end, which reformulation is worth trying. If routine mathematical work becomes automated, fewer people may have a reason to pass through the apprenticeship that builds this knowledge. The risk is not that papers vanish. The risk is that the papers remain while fewer people know how to use them well. This section should stay concrete. The issue is cumulative knowledge, not nostalgia for hand calculation.

The education section follows naturally. The same AI tool can amplify experts and weaken novices because the expert already has internal checks. An expert can reject a bad proof sketch, notice a missing hypothesis, or see that an argument proves the wrong thing. A novice may receive the same output while bypassing the attempts and repairs that would have built those checks. This is a training problem rather than a cheating story. If execution becomes cheap, the curriculum has to preserve the activities that form judgment: checking, posing questions, comparing approaches, explaining failures, and deciding which results matter. The sharper question is what training produces people who can direct more capable tools later.

The institutional turn comes after education because institutions shape what gets practiced. Benchmarks and verifiers begin as useful proxies: they make progress visible. Over time, visible progress can become the thing that gets funded, hired for, and trained toward. AI strengthens this loop because it produces gains precisely where the proxy is sharp. This is the measurability distortion: tractable work starts to look like important work because it is the part that moves on the scoreboard. The article does not need a broad attack on benchmarks or institutions. The mechanism is enough. When a community repeatedly rewards what is easy to score, young researchers learn the taste of that reward system.

The conjecture section turns the institutional point back toward mathematics. If proof becomes easier to obtain, good questions become relatively more valuable. The hard work moves upstream: choosing the problem, finding the right definition, deciding which analogy is worth developing, seeing why a statement should organize later work. The Erdos example is useful if handled carefully. If AI proves many Erdos-style conjectures, the achievement of asking those questions becomes more visible. The proof closes the problem; the conjecture selected a problem worth closing. The claim is comparative: when one layer becomes cheaper, value shifts toward the layers that remain hard to score and hard to automate.

Only after this does the article generalize to science. Fields are too coarse as prediction units. Biology contains fast-feedback structure prediction and slow-feedback clinical intervention. Mathematics contains formal proof checking and open-ended conjecture formation. Physics contains simulation-tight subproblems and theory choices that may wait years for evidence. Economics and climate contain questions where controlled feedback is slow, noisy, or impossible. The feedback-cost map asks how quickly candidate answers become reliable training or evaluation signal. Where that loop closes, AI can compound. Where it does not, progress may still happen, but the mechanism is different.

The ending uses this map to handle plateau, AGI, and labor without making the article too broad. There is no single plateau question: progress can continue where checks are cheap while slowing where feedback is absent, delayed, or expensive. For this article, the useful question is not whether a system deserves a general label, but which feedback loops it can close. The same split reframes labor. If the human in an AI-assisted workflow is only a serial review bottleneck, there is economic pressure to reduce that bottleneck. But roles tied to question selection, responsibility, interpretation, and standards sit partly outside the cheap-feedback pipeline. The article ends by returning to the opening test: AI already does parts of mathematics and science; the question is which feedback loops it can close, and what happens to the valuable work those loops do not measure.