Paper

A benchmark of nine factor-mining methods finds no approach consistently wins

FactorBench compares roughly five thousand machine-mined factors across five equity markets, from genetic programming to LLM agents.

AI Fin ResearchCovers Five equity marketsWorking paper, not peer reviewed

Zhuohan Wang and Carmine Ventre introduce FactorBench, a benchmark for automated factor mining. It compares roughly five thousand mined factors from nine automated methods across five equity markets. The methods span genetic programming, reinforcement learning, generative models and large language model agents.

A shared data and evaluation contract accepts both symbolic expressions and executable Python factors, so different discovery algorithms feed the same signal-combination and portfolio-construction steps. The benchmark looks at three levels. The first is factor validity, temporal generalization and predictiveness beyond measured risk and style exposures. The second is how distinct the factor pools are within and across methods, including similarity to the Alpha101 set. The third is composite-signal quality and after-cost long-only and long-short portfolio performance.

The finding is that no paradigm consistently dominates. The abstract does not name the five markets or the sample period.

Sources

Related