The Factor Zoo: Machine Learning, Multiple Testing, and Return Predictability

Authors

  • Mengqi He Anhui Polytechnic University, China

Keywords:

Factor Zoo, Multiple Testing, Data Snooping, Machine Learning, Asset Pricing, Return Predictability, Harvey, Replication Crisis

Abstract

The asset pricing ‘factor zoo’—the proliferation of hundreds of purportedly significant stock return predictors published in academic finance journals—has created a credibility crisis for empirical asset pricing. Harvey, Liu, and Zhu (2016) in the Review of Financial Studies documented that over 300 factors had been published by 2014, raising serious concerns about data snooping (the inflation of apparent statistical significance from testing many hypotheses on overlapping datasets) and multiple testing (the expected false discovery rate when many hypotheses are tested without adjustment). The typical t-statistic threshold of 2.0 used to claim statistical significance is woefully inadequate when hundreds of factors are tested: at t > 2.0, 5% of true null hypotheses are rejected by chance, and with 300 factors this implies 15 spurious ‘discoveries.’ McLean and Pontiff (2016) document that published factor returns decay substantially after publication—suggesting many factors are overfit in-sample. This paper reviews the factor proliferation problem, multiple testing corrections, machine learning approaches to factor discovery, and the standards required for claiming valid return predictability.

Downloads

Published

2026-06-01