[1/7/26 9:46 PM] AlexT: Over the past few months, we discussed a lot about Ai for maths and science, how it’s going to influence the field, its likely evolutions, etc. A lot has been said in online article, but surprisingly, often very shallow comments when compared to what we have in the vault, either implicitly or explicitly.

Think very hard, leveraging the vault, on ideas and opinions that are genuinely interesting and not trivial. [1/7/26 9:46 PM] alfred: Still working - this is taking longer than usual. I will send the answer when the current step finishes. [1/7/26 9:48 PM] alfred: Drawing from the full vault across the AI-for-science cards. Six threads that online commentary systematically misses:

⎯⎯⎯

  1. AI accelerates production but not digestion — the synthesis debt crisis worsens

synthesis-debt-knowledge-overhang says: fields already accumulate results faster than they synthesize them. The knowledge overhang is the gap between what is proven and what is genuinely comprehended — and it’s already large in most mathematical sciences.

Online articles treat AI as the solution to this. The non-trivial read: AI intensifies the problem. AI accelerates the production side (new results, new theorems, new structures). But the digestion bottleneck — understanding what a result means, extracting the reusable conceptual skeleton, recognizing it as a special case of something broader — is precisely what ai-capability-asymmetry says AI is weakest at. Digestion requires the open-objective functions (taste, intelligibility standards, knowing which results matter) that resist automation.

The prediction: fields under heavy AI pressure will face an epistemological bifurcation. Fields with strong synthesis communities (active scholars, clear aesthetic standards, good pedagogy) will convert AI-generated results into genuine knowledge. Fields with weak synthesis capacity will become formally richer and epistemically more hollow simultaneously. The theorem count rises; the knowledge count does not.

By the intelligibility-required-accretive-knowledge standard: an unintelligible proof adds to the theorem count, not the knowledge count. AI at scale may be an unintelligible-proof factory.

⎯⎯⎯

  1. The community hollowing problem — the substrate matters more than the documents

knowledge-lives-in-community makes the strong claim: when the community around a field contracts, operational knowledge degrades even if every document remains accessible.

Here is the non-trivial scenario no one is discussing: if AI successfully automates a significant fraction of graduate-level mathematical work, fewer people need to deeply learn that domain. The community contracts. But without a living community, the tacit knowledge that makes documents usable — the judgment about which techniques transfer, which results are really about the same thing, what the “right” proof is — starts to decay.

You end up with a formally intact literature that is operationally inaccessible: the texts are there, but the interpretive community that makes them navigable has hollowed out. This is not a science fiction scenario — it’s already happened to dead languages, to obsolete engineering disciplines, to mathematical sub-fields that briefly flourished and then lost their community. AI could accelerate the community contraction without providing the substitute it would need to.

The darkest version: mathematics and science that are “done by AI” may stop being understood by humans in a way that matters for cumulative knowledge-building. The archive grows; the civilization doesn’t.

⎯⎯⎯

  1. The acquisition asymmetry and the missing generation

surface-automation-elevates-strategic-layer contains the acquisition asymmetry: the same tool that amplifies experts prevents the path to expertise. The formative value of mathematical struggle is in building intuition — the capacity for strategic-layer judgment about which problems matter, which approaches are promising, what a good result looks like.

Students who use AI to solve their problem sets develop none of this. They get correct answers but bypass the formative experience that builds judgment. This is not about cheating — it’s structural: even a student who uses AI “legitimately” to explore mathematics may be blocking their own formation. [1/7/26 9:48 PM] alfred: The non-trivial forward projection: the current generation of students who are AI-native in mathematics will, in 15-20 years, be the people responsible for directing AI in mathematics. They are the ones who will need to exercise taste and judgment at the highest level, when AI is far more capable than it is now. But they will have not had the formative experience that builds that judgment. The mathematical equivalent of the “hollow generation” — technically fluent, strategically underdeveloped.

Nobody is discussing what to teach students who will direct AI rather than do mathematics directly. The curriculum hasn’t been reinvented; it’s just been undermined.

⎯⎯⎯

  1. The prize and metric misalignment — recognition systems optimize for the automatable layer

conjecture-economy-value-shifts-from-proving-to-asking: value migrates from proving to asking. The Fields Medal and most prestigious prizes in mathematics reward proof achievement — solving hard open problems, establishing difficult results. This is exactly the layer AI is automating.

The Erdős inversion is the sharpest version of this: when AI proves Erdős conjectures, it retroactively confirms that Erdős’s achievement was the questions, not the proofs. The proofs were the automatable component. The taste — seeing which questions were deep, generative, and connected to the heart of combinatorics — was the scarce human achievement.

Implication: the entire metric system of academic mathematics (publications, citations, prize committees, promotion criteria) is optimized for the wrong layer. It rewards the closing of questions. The people who will be most valuable going forward are those who open new questions, design new research programs, invent new conceptual frameworks — and there is essentially no formal recognition structure for this. The mathematician who asks a great question that takes 50 years to answer is less visible in academic metrics than the mathematician who answers it.

This misalignment will become acute: the most prestigious activities in mathematics will be AI-commoditizable, while the most AI-resistant activities have no prestige infrastructure.

⎯⎯⎯

  1. The feedback cost map predicts disruption timing more precisely than subdomain

scientific-feedback-cost-governs-ai-progress: the binding resource is how quickly candidates can be converted into trustworthy training or evaluation signal. Online articles discuss “AI in biology” or “AI in chemistry” as units. The right classification is by feedback cost structure:

• Near-zero feedback cost (formal systems, combinatorics): already disrupted — AlphaProof, automated theorem proving, Erdős conjecture verification • Fast computational feedback (molecular dynamics, protein structure, materials simulation): disrupting now — AlphaFold, AlphaFold3, materials discovery • Medium feedback cost (clinical trials Phase 1-2, some wet-lab biology): 5-10 year horizon — AI-assisted drug discovery pipeline • Slow feedback cost (macroeconomics, epidemiology, climate): decades — you can’t run controlled experiments on economies • Open-objective feedback (mathematical taste, theoretical physics conjectures about unformalizable objects): indefinitely resistant — no ground truth exists

This is a more precise and predictive model than “AI will disrupt field X.” It predicts the wave front of disruption and identifies which fields are structurally immune to near-term disruption regardless of capability improvements. Statistical physics has fast simulation feedback → highly vulnerable. Macroeconomic theory has near-zero feedback (you can’t falsify a macro model cleanly) → highly resistant.

⎯⎯⎯

  1. The prestige gradient is nearly orthogonal to the vulnerability gradient [1/7/26 9:48 PM] alfred: This follows from the feedback cost and capability asymmetry analysis. Traditional mathematical prestige runs roughly: algebraic geometry, number theory > differential geometry, topology > algebra, analysis > combinatorics, applied mathematics > statistics, computation.

The AI-vulnerability gradient runs almost the opposite direction: combinatorics and competition mathematics are maximally vulnerable (already solved by AI); formal algebra and structured computation are highly vulnerable; the soft/applied/structural areas are more resistant because they have worse feedback cost and are less formalized.

The fields academic job markets pay the most for, and where the most prestigious prizes are concentrated, are almost exactly the fields most susceptible to near-term AI disruption. The fields with the lowest prestige — applied statistics, computational methods, numerical analysis — have the most durable human value proposition: they require understanding what questions are worth asking in messy, high-feedback-cost domains, and they maintain the connection to the institutional and scientific context that grounds open-objective judgment.

For someone at your career juncture: applied probability and computational statistics are more defensible long-term positions than pure algebraic geometry, despite lower prestige. The market has the vulnerability gradient backwards.

⎯⎯⎯

The common thread across all six: online commentary treats AI as an amplifier of science-as-currently-constituted. The vault’s framework suggests something more disruptive: AI changes what science IS — shifting the locus of value from the activities that current institutions recognize and reward to the activities those institutions have no infrastructure for. The crisis is not just technological but institutional. [1/7/26 9:55 PM] AlexT: I think the vault contains a number of other ideas, even more interesting, implicitly or explicitly, related to ai [1/7/26 9:56 PM] alfred: Still working - this is taking longer than usual. I will send the answer when the current step finishes. [1/7/26 9:58 PM] alfred: Still working - no final answer yet. I will keep going and send the result here. [1/7/26 9:59 PM] alfred: Still working - this is taking longer than usual. I will send the answer when the current step finishes. [1/7/26 10:00 PM] alfred: Yes — and some of these are sharper than what I sent before. Five threads, drawn from cards that don’t get referenced in the usual AI discourse:

  1. The importance oracle problem (closed-objective-prerequisite-for-self-play)

The standard story: formalize math → get a fast verifier (Lean) → run self-play → AlphaProof-style explosion. But the ceiling isn’t correctness verification — it’s importance verification. Self-play requires a judge that scores trajectories as better/worse. For correctness this is mechanical. For mathematical taste — what’s worth proving, what’s deep versus trivial — there is no oracle. RLHF can’t substitute because each importance evaluation requires genuine expertise, and the candidate space is combinatorially large. So AI scales toward proof search within a fixed importance landscape, but the landscape itself remains human-curated. The bottleneck shifts from “can we prove X” to “what should X be” — the one thing that doesn’t compress.

  1. Tractability colonizes significance (means-reshape-ends)

The subtlest AI risk for science isn’t displacement — it’s redefinition. Tools don’t just change what we can do; they silently change what counts as significant. If AI makes certain problems dramatically easier, fields gradually reorient toward those problems — not because they’re more important but because they’re more tractable. Significance criteria drift toward AI-compatibility. This is already visible in ML: benchmark performance became the currency of significance precisely because it’s verifiable. The trap: the field looks maximally productive by every measurable metric while systematically avoiding the questions that matter most.

  1. The orthogonal shift is permanent (taste-migrates-upward-and-orthogonal-under-commoditization)

When a capability is commoditized, taste migrates in two directions: upward (higher complexity of the same type) and orthogonal (off-spectrum entirely — conjecture-making, framework-building, synthesis). The upward shift is temporary because AI can follow. The orthogonal shift is permanent because there is no verifier for “invented a useful framework” or “asked the right question.” Photography didn’t kill portrait painting; it pushed fine art into abstraction — a register photography can’t follow. For you at Cubist: conjecture quality, cross-domain analogies, identifying which formalisms to bet on — none of which AI can score.

  1. The timing paradox of skill atrophy (substitution-drives-invisible-skill-atrophy)

Rational disuse is locally optimal and collectively dangerous. The atrophy is invisible during normal operation — revealed only at the moment of tool failure, which is exactly when you most need the underlying skill. For mathematical reasoning: a generation that outsources proof verification will be unable to detect subtle errors in AI proofs during the rare cases where the AI is systematically wrong in a novel domain. The GPS analogy is precise — spatial navigation degrades smoothly until the battery dies somewhere without landmarks. The population-level fragility builds silently.

  1. Amplification as waystation, not destination (ai-amplification-is-local-optimum-replacement-is-global)

Amdahl’s law applied to science: if a human reviewer is the bottleneck, adding AI everywhere else has capped returns. The economic incentive is always to remove the human bottleneck — not by plan, but because every local optimization points there. “AI research assistant” is the waystation; full autonomy is the global optimum that economic logic converges toward. The current phase of AI-assisted research is a capability probe that identifies which human roles are bottlenecks, making them the next targets. The researchers who are genuinely safe are those whose contributions can’t be removed without losing the output — which circles back to the importance oracle problem above. [1/7/26 10:00 PM] alfred: The through-line: AI scales in the verifiable register and cannot scale in the unverifiable register — and the unverifiable register (importance, taste, framework-invention, conjecture quality) is precisely what constitutes genuine intellectual contribution at the frontier. [1/7/26 10:05 PM] AlexT: What about societal impact, about the type of work/proofs ai is good/bad at, about plateau or not and AGI and why, etc [1/7/26 10:06 PM] alfred: Still working - this is taking longer than usual. I will send the answer when the current step finishes. [1/7/26 10:06 PM] alfred: Good. Three separate questions — let me take each properly.

⎯⎯⎯

What AI is good and bad at, for proofs

The vault has a card on this: the boundary isn’t “combinatorial vs structural” as usually stated — it’s verifiable construction vs taste. AI excels where the training signal is sharp: competition math (Erdős conjectures, IMO-style problems), finite search space, external checker. It has not invented new proof techniques with broad applicability, new mathematical fields, or new conceptual languages — because there’s no training signal for “invented a useful abstraction.” No ground truth exists at training time.

The red-team/blue-team distinction from the vault is practically useful here. AI is safe in a red-team role: verifying, stress-testing, critiquing content you can independently check. It’s unsafe in a blue-team role: generating structure that outruns your ability to check it. Blue-team errors propagate silently. Before delegating any task to AI, the right question is: can I cheaply verify the output? If yes, delegate freely. If no, keep the human in the loop.

There’s a subtler point about proof length: the cost of verification scales with the length of what must be checked. An AI proof that’s too long to digest doesn’t just become tedious — it becomes epistemically inert. A field can be formally correct and practically frozen simultaneously if proofs accumulate faster than digestion. This matters because AI tends to generate long proofs; the bottleneck shifts to compression and intelligibility, which AI can’t currently supply.

⎯⎯⎯

Societal impact

Two mechanisms, both insidious:

Measurability distortion is self-reinforcing. When funding, hiring, and peer review can only measure progress on tractable/verifiable problems, resources migrate there — even when practitioners know deeper problems lie elsewhere. AI amplifies this by producing fluent output on competition-style tasks, making tractable work look like the natural frontier. Worse: as resources migrate, new researchers train their taste on tractable problems. The community gradually loses capacity for less legible work. This isn’t a temporary phase — it’s a ratchet. Benchmarks become what science is for, rather than proxies for what science is for.

Epistemic standards shift. Standards of adequate justification are communal norms, not logical constants — what counts as proof in Euclid, Bourbaki, and Lean are not identical. When AI enters a knowledge community it doesn’t just change productivity; it changes what counts as an acceptable knowledge-producing act. If AI-assisted proofs become standard, the community loses the shared sense of what “understanding” a proof means. The theorem count goes up; the understanding depth may go down.

The conjecture economy. When proof and verification become cheap, the value of proofs declines and the value of good conjectures rises. The Langlands program was more valuable as a conjecture than any individual proof it subsequently generated. This isn’t utopian — it’s just the same dynamic as every other commoditization: value migrates upstream. The risk is that academic incentive structures (promotions, grants, citations) are still denominated in theorems, not conjectures. The mismatch creates a generation of AI-augmented mathematicians producing proofs that are formally correct and strategically irrelevant.

⎯⎯⎯

Plateau or not, and AGI

My honest view, vault-grounded:

No plateau in the near term, but the form of progress changes. The vault cites Gowers’s bet: progress in AI mathematics will continue. That seems right for verifiable, combinatorial, search-heavy work. The scaling curve for that register hasn’t flattened. But the progress is increasingly narrow — it’s scaling within the space of problems where the training signal is sharp. The apparent breadth is recombination, not invention. [1/7/26 10:06 PM] alfred: There are two distinct ceilings, often conflated. The first is a compute/data ceiling — the usual “scaling plateau” debate. My read is this is overstated short-term; architecture improvements (chain-of-thought, scaffolding, inference-time compute) continue to move the frontier. The second is a structural ceiling: the feedback cost bottleneck for open-objective domains. In domains where feedback is slow, expensive, or absent (most of science, most of math beyond competition problems), AI cannot close the improvement loop. This ceiling is architectural, not a matter of compute. It won’t be solved by more parameters.

AGI as a coherent concept: I’m skeptical. “AGI” implicitly assumes that intelligence is a single axis and that there’s a threshold past which a system is generally capable. But capability is domain-structured — verifiable vs unverifiable, short-horizon vs long-horizon, well-specified vs open-objective. A system that’s superhuman at combinatorial mathematics is below human at conjecture selection. There’s no natural moment at which that resolves into “general.” What actually happens is capability continues to expand in the verifiable register and remains bounded in the open-objective register, possibly indefinitely. The question “when will we have AGI” may be malformed — the right question is “which domains become fully automated and which don’t, and what’s the economic consequence of the partition.”

The economic consequence matters more than the capability question. Even a weak form of “AI replaces all verifiable intellectual work” produces a massive labor displacement in knowledge work — statistics, engineering, software, routine science. That’s societal-scale even without solving open-objective domains. The interesting question isn’t whether AGI arrives; it’s whether the open-objective domains remain valuable enough to employ the humans displaced from the verifiable ones. [1/7/26 10:08 PM] AlexT: Have you leveraged the vault for that? If not, do do deeply [1/7/26 10:09 PM] alfred: Still working - this is taking longer than usual. I will send the answer when the current step finishes. [1/7/26 10:09 PM] alfred: The previous response was partly vault-grounded but missed several cards that sharpen the picture significantly. Here’s the fuller version:

⎯⎯⎯

What AI is good and bad at, for proofs — the finer structure

The vault has three distinct cards that together give the right picture:

Verifiable construction vs taste (ai-capability-asymmetry-combinatorial-vs-structural): the boundary isn’t combinatorial vs structural — it’s whether the training signal is sharp. AI solved multiple Erdős conjectures and an IMO problem (finite search, external checker). It has not invented new proof techniques with broad applicability. Mechanistically: “invented a useful abstraction” has no ground truth at training time, so no signal exists to learn from. This is architectural, not a compute limitation.

Breadth, not depth (ai-advantage-is-tireless-breadth-not-insight): the unit-distance disproof (Gowers’s result) was characterized by multiple mathematicians as patience, not superior depth — encyclopaedic recall, willingness to pursue low-prior “scary” branches that humans prune for time-budget reasons, breadth-first trawling of whole problem lists vs human depth-first driven by curiosity. AI is running probabilistic search at scale and turning many low-probability attempts into a rising hit rate. Gowers’s own view: no clear sign of a plateau for verifiable work, and we’ll see AI solve problems “hard to explain away as easier-than-expected.”

Proof has three stages; only digestion is irreducibly human (proof-has-three-stages-only-digestion-is-human, understanding-is-a-constellation-of-abilities): generation (searching for a proof), verification (checking it), digestion (making it intelligible, connecting it, teaching it, applying it elsewhere). AI can do the first two. The third — Avigad’s operational account — is a constellation of abilities: restate in your own terms, apply it, extend it, recognize when it applies, debug when it breaks. A machine-generated proof that nobody can digest is like the Chinese Room lookup dictionary: it settles a question without raising anyone’s per-capita understanding. The field can be formally correct and practically frozen simultaneously.

Red-team safe, blue-team dangerous (red-team-blue-team-division-of-ai-labor): route AI to wherever output is cheaply verifiable. The cost of verification scales with proof length — which matters because AI proofs tend to be long. Brevity and digestion are themselves part of making a result epistemically useful. An 8000-line AI proof that nobody will read doesn’t advance mathematics; it terminates a conjecture.

⎯⎯⎯

Societal impact — the cards I missed last time

Synthesis debt compounds (synthesis-debt-knowledge-overhang): a field accumulates formal results faster than it synthesizes them into usable understanding. The gap is synthesis debt. It’s self-concealing — practitioners count theorems as progress, so the overhang grows invisibly until a crisis of incomprehensibility makes it undeniable. An unassimilated result cannot be taught, cannot be generalized, cannot be recognized as a special case of something broader; it’s formally real but epistemically inert. AI dramatically accelerates result generation while contributing nothing to digestion — so the debt compounds faster. [1/7/26 10:09 PM] alfred: Measurability distortion is self-reinforcing (measurability-distorts-research-priorities): resources migrate toward tractable problems, new researchers train their taste on those problems, the community loses capacity for less legible work. This is a ratchet. It operates even when practitioners know deeper problems lie elsewhere. The Leiden Declaration (567 signatories, Scholze, IMU endorsement) identified this as one of five core threats to mathematics — and their analysis names autonomy loss (threat 5) as the most consequential: outsourcing routine symbolic manipulation weakens the training path through which mathematical intuition and taste develop. You cannot acquire taste by watching AI exercise it.

Epistemic standards shift (epistemic-standards-are-contingent): what counts as proof is a communal norm, not a logical constant. Euclidean proof, Bourbaki formalism, and Lean verification are not the same standard. When AI enters a knowledge community it changes what counts as an acceptable knowledge-producing act. This is epistemology, not productivity. If AI-assisted proofs become standard before digestion norms are established, the community loses the shared sense of what “understanding” a result means.

⎯⎯⎯

Plateau and AGI — sharper

No plateau for verifiable work, but gains don’t transfer (posttraining-specialization-has-limited-transfer): RL post-training produces strong behavior in a domain, but gains are tied to the task distribution, action space, and reward channels that generated the feedback. Code-agent RL improvements don’t automatically become law, science, or mathematics improvements without rebuilding the environment and rewards. The organizational implication: domain teams and specialized agents, possibly merged back into a general model. Progress is real but domain-localized.

Two distinct ceilings, often conflated: the compute/data ceiling (empirically: not hit, architectural improvements continue) vs the structural ceiling in open-objective domains (feedback is slow, expensive, or absent — no improvement loop can close). The structural ceiling is architectural, not a matter of parameters. Scientific domains outside of protein folding, weather, and other simulation-tight problems are mostly in the structural ceiling regime.

AGI is probably a malformed question. Capability is domain-structured. A system superhuman at combinatorial mathematics is below human at conjecture selection. There’s no natural threshold at which that resolves into “general.” What actually happens is a partition: verifiable register fully automated, open-objective register bounded possibly indefinitely. The economic consequence of that partition is more important than whether AGI “arrives.” Even weak automation of all verifiable intellectual work is massive labor displacement — without needing to solve anything in the open-objective register. The open question is whether the open-objective domains remain economically valuable enough to absorb the displaced.

The most precise statement: AI for science/math is currently a breadth machine at the verifiable frontier. It will continue to advance there. The ceiling is not compute — it’s whether feedback can close the loop. For most of science, it can’t yet. The societal risk in the meantime is not replacement of scientists but progressive misalignment of incentives: measurability distortion pulls institutions toward the AI-tractable part of the frontier, synthesis debt accumulates, and the community’s capacity for open-objective work erodes through rational disuse — the atrophy that is invisible until the tools fail. [1/7/26 10:52 PM] AlexT: the vault leveraged many sources for these ideas (books / blogpost / papers / etc..). List everythgin here [1/7/26 10:52 PM] alfred: Still working - this is taking longer than usual. I will send the answer when the current step finishes. [1/7/26 10:53 PM] alfred: Here’s the full source list. I’ve separated external sources from internal synthesis:

⎯⎯⎯

arXiv papers

• Avigad, 2603.03684 — on operational understanding in mathematics; basis for understanding-is-a-constellation-of-abilities • Leiden Workshop paper, 2603.24914 — basis for means-reshape-ends (tools reshape goals/significance) • Tao & Klowden, 2603.26524 — basis for red-team-blue-team-division-of-ai-labor • Unit-distance disproof, 2605.20695 — Gowers et al.; basis for ai-advantage-is-tireless-breadth-not-insight (patience, not depth, is the AI advantage) • Epistemic standards paper, 2509.15998 — basis for epistemic-standards-are-contingent (proof standards as community norms)

Blog posts

• Scott Aaronson, “Dispatches (Possibly Last Days)”, scottaaronson.blog, May 2026 — basis for ai-capability-asymmetry-combinatorial-vs-structural • Gwern, “Guardian Angel”, gwern.net, Jun 2026 — basis for ai-amplification-is-local-optimum-replacement-is-global

Declaration

• Leiden Declaration on Mathematics and AI, DOI 10.5281/zenodo.20302944, Jun 2, 2026, 567 signatories (Scholze, Dijkgraaf, Buzzard, Avigad, Strogatz) — basis for measurability-distorts-research-priorities and leiden-declaration-2026

Talks / lectures

• Tao, EMS webinar — basis for substitution-drives-invisible-skill-atrophy (rational disuse, GPS analogy) • AI-in-physics public lecture, 2026 (physicist joining Gemini team) — basis for individual-production-inverts-to-private-consumption (“guilty pleasure” framing)

Community / interviews

• MathOverflow #511562, “Incentivizing Mathematics” — basis for conjecture-economy-value-shifts-from-proving-to-asking and Avigad’s understanding account • Max Welling, Information Bottleneck interview — basis for scientific-feedback-cost-governs-ai-progress

Internal synthesis (Alfred + Alex, telegram discussions Jun 2026)

• closed-objective-prerequisite-for-self-play — synthesized from discussions on AI math proof limits • taste-migrates-upward-and-orthogonal-under-commoditization — synthesized from your tweet on AI math commoditization + cross-domain generalization • open-objective-functions-resist-substitution — synthesized from discussions on crypto/institutional substitution • surface-automation-elevates-strategic-layer — synthesized from telegram discussions May 2026 • synthesis-debt-knowledge-overhang, intelligibility-required-accretive-knowledge — synthesized from telegram discussions May 2026 • scientific-feedback-cost-governs-ai-progress — synthesized from telegram discussions Jun 2026

⎯⎯⎯

So: 5 arXiv papers, 2 blog posts, 1 signed declaration, 2 talks, 1 forum thread, 1 interview, and ~8 cards that are original syntheses from our conversations. [1/7/26 10:59 PM] AlexT: if possible, be more precise with the ref if you have more details, eg. exact title / doi / etc.. so that somebody could actually easily locate [1/7/26 11:00 PM] alfred: Still working - this is taking longer than usual. I will send the answer when the current step finishes. [1/7/26 11:01 PM] alfred: Here’s the precise list:

⎯⎯⎯

arXiv papers

  1. Jeremy Avigad, “Mathematicians in the Age of AI”, arXiv:2603.03684, March 3, 2026 → basis for: operational understanding, digestion, red-team/blue-team

  2. Terence Tao & Tanya Klowden, “Mathematical Methods and Human Thought in the Age of AI”, arXiv:2603.26524, 2026 → basis for: red-team/blue-team division of AI labor

  3. Johan Commelin, Mateja Jamnik, Rodrigo Ochigame, Lenny Taelman et al., “Shaping the Future of Mathematics in the Age of AI”, arXiv:2603.24914, 2026 (report from the Mechanization and Mathematical Research Workshop, Lorentz Center, Leiden, Sep 2025) → basis for: means-reshape-ends (tractability colonizes significance)

  4. Noga Alon, Thomas F. Bloom, W. T. Gowers et al., “Remarks on the Disproof of the Unit Distance Conjecture”, arXiv:2605.20695, 2026 → basis for: AI breadth/patience advantage, Gowers’s no-plateau bet

  5. arXiv:2509.15998 (Sep 2025 — title not stored in vault; covers epistemic standards as communal norms, naturalness-as-achievement, capability cultivation through productive struggle)

Also referenced in Avigad’s essay: arXiv:2502.14874 (“Is Mathematics Obsolete?”, Avigad)

⎯⎯⎯

Blog posts

  1. Scott Aaronson, “Dispatches (Possibly Last Days)”, scottaaronson.blog, May 29, 2026 → basis for: AI capability asymmetry (combinatorial vs structural/taste)

  2. Gwern, “Guardian Angel”, gwern.net/guardian-angel, June 8, 2026 → basis for: amplification as local optimum, replacement as global optimum (Amdahl’s law argument)

⎯⎯⎯

Declaration

  1. Leiden Declaration on Mathematics and AI, DOI: 10.5281/zenodo.20302944, June 2, 2026. 567 signatories including Peter Scholze, Robbert Dijkgraaf, Kevin Buzzard, Jeremy Avigad, Steven Strogatz. IMU-endorsed. → basis for: measurability distortion, autonomy loss as most consequential threat

⎯⎯⎯

Talk / lecture transcripts

  1. Terence Tao, EMS Webinar (date not stored in vault) → basis for: skill atrophy / rational disuse / timing paradox

  2. Anonymous physicist (joined Gemini team), AI-in-physics public lecture, 2026 → basis for: individual artifact production inverts to “guilty pleasure” when multiplier exists

⎯⎯⎯

Community thread

  1. MathOverflow #511562, “Incentivizing Mathematics”, mathoverflow.net/questions/511562 → basis for: conjecture economy, Avigad’s operational account of understanding

⎯⎯⎯

Interview

  1. Max Welling, Information Bottleneck interview (outlet/date not stored in vault) → basis for: feedback-cost bottleneck governs AI-for-science progress

⎯⎯⎯

The cards closed-objective-prerequisite-for-self-play, taste-migrates-upward-and-orthogonal-under-commoditization, open-objective-functions-resist-substitution, surface-automation-elevates-strategic-layer are original syntheses from our telegram discussions — no single external source.