Paper

In 912 AI-written economics papers, the ideas trail human work far more than the execution

Against 41 papers from the AER and AEJ: Economic Policy, a study puts 71% of the quality gap on the research idea and finds no significant gap on robustness.

AI Fin ResearchCovers GlobalWorking paper, not peer reviewed

Ning Li compares 912 economics papers generated entirely by AI in the APE project with 41 human papers published in the American Economic Review and AEJ: Economic Policy. The paper splits quality in two. Idea quality is scored by a pair of language models fine-tuned on publication decisions. Execution quality is scored on a six-part rubric by Gemini 3.1 Flash Lite, the model family that judges the APE tournament.

Human papers AI papers Gap (Cohen’s d)
Idea quality, as the mean probability of an exceptional idea 47.1% 16.5% 2.23
Execution quality, out of 5 4.38 3.84 0.90

Idea quality accounts for about 71% of the overall difference and execution for 29%. Within execution, the largest weakness is the depth of mechanism analysis (d = 1.43). The study finds no significant difference on robustness.

The AI papers lean on one design: 74% use difference-in-differences. Seven of the 912, or 0.8%, score above the median human paper on both ideas and execution.

Both measures come from language models and not from human referees. The human sample is 41 papers from two journals.

Sources

Related

Analysis

Throw out what you knew about AI before September

Fable 5.1, GPT-6 Astra and Opus 5.5 arrived inside five weeks. One science benchmark doubled in two months. A test you ran in the spring describes a different technology.

Anthropic, Introducing Claude Opus 5.5Global