Paper

LLM literature reviews hold up for broad claims and break down for paper-level ones

A test on economics papers that use rainfall as an instrument finds accuracy falls as the reading task needs more context.

AI Fin ResearchWorking paper, not peer reviewed

Jeffrey D. Michler, Kieran Douglas and Anna Josephson treat the LLM-assisted literature review as a measurement problem. They use three implementations of ChatGPT to identify economics papers that use rainfall as an instrumental variable and to extract metadata from them. Each implementation is benchmarked against a human-labeled subset and then run on the full corpus.

The models do well on binary classification. Performance deteriorates as tasks demand more contextual interpretation. The main result is about which claims the extracted data can support: the same amount of measurement error substantially affects paper-level claims and has little effect on broader claims about the literature. The authors point out that this puts the error at the level of detail where an LLM adds the most over a human reviewer.

They conclude that standard performance metrics describe the quality of generated data but do not by themselves make downstream inference credible. The abstract reports no headline accuracy figures, and the test covers one model family and one literature.

Sources

Related