







I think there should be a norm that when a set of AI solutions to mathematical problems is released, the set of all problems attempted be released alongside.When trying to understand AI capabilities, it is problematic to selectively report only positive results. https://t.co/ywcWCnFojS— Itai Sher (@itaisher) August 1, 2026
Mathematicians are grappling with the possibility that AI might eclipse them
I talked to 20 mathematicians about rapid AI progress in their field.

The AI Question that No AI Person Asks
The AI Question that No AI Person Asks
The fall of the theorem economy
How AI could destroy mathematics and barely touch it

The fall of the theorem economy
How AI could destroy mathematics and barely touch it

The AI Revolution in Math Has Arrived | Quanta Magazine
AI is being used to prove new results at a rapid pace. Mathematicians think this is just the beginning.


Knowledge Collapse
AI companies are racing to mechanize mathematics. Where does that leave human understanding?

Daron Acemoglu on Twitter / X
I recommend Columbia mathematician Michael Harris’s wide-ranging, informative and thought-provoking essay in Boston Review on AI and mathematics:https://t.co/txwAd8ri4xHarris rightly worries about the possible negative effects of AI-generated proofs and mathematics on…— Daron Acemoglu (@DAcemogluMIT) June 16, 2026
OpenAI’s math breakthrough played to AI’s strengths
I tried to explain OpenAI’s solution more clearly than OpenAI did.

Mathematical methods and human thought in the age of AI
Artificial intelligence (AI) is the name popularly given to a broad spectrum of computer tools designed to perform increasingly complex cognitive tasks, including many that used to solely be the...

What it Means to Be a Mathematician When AI Does the Math
Researchers debate motivation, purpose, and the field’s future

The argument against AI agents and unnecessary automation
Opinion: OpenAI's Operator a solution in search of a problem

📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani
"AI" is Automated Inequality
Tech bros still dominate the discussions about so-called "AI" with false claims. Even most "AI"-critical researchers spend much of their time meticulously debunking (always only a subset of) claims, leaving vast areas of the economic consequences of "AI" unexplored. (Even the "AI"-evangelist Economi

State of AI Report 2025
The State of AI Report analyses the most interesting developments in AI. Read and download here.
