[Paper Review] Missing Values Handling for Machine Learning Portfolios
The paper analyzes missingness structure in 159 cross-sectional return predictors and shows that simple cross-sectional mean imputation often performs as well as EM-based imputations for machine learning portfolios. It explains why observed data offer limited information about missing values due to time-structured blocks and weak cross-sectional correlations.
We characterize the structure and origins of missingness for 159 cross-sectional return predictors and study missing value handling for portfolios constructed using machine learning. Simply imputing with cross-sectional means performs well compared to rigorous expectation-maximization methods. This stems from three facts about predictor data: (1) missingness occurs in large blocks organized by time, (2) cross-sectional correlations are small, and (3) missingness tends to occur in blocks organized by the underlying data source. As a result, observed data provide little information about missing data. Sophisticated imputations introduce estimation noise that can lead to underperformance if machine learning is not carefully applied.
Motivation & Objective
- Characterize the structure and origins of missingness for 159 cross-sectional return predictors.
- Evaluate missing value handling strategies for portfolios constructed with machine learning.
- Assess when sophisticated imputations may add or hinder predictive performance in portfolio contexts.
Proposed method
- Characterize missingness patterns and origins in predictor data.
- Compare simple cross-sectional mean imputation to expectation-maximization imputations.
- Assess impact of imputation on machine learning-driven portfolio construction.
Experimental results
Research questions
- RQ1What is the structure and origin of missingness in the predictor data?
- RQ2How does cross-sectional mean imputation compare to EM methods for portfolio construction using ML?
- RQ3Under what conditions do sophisticated imputations improve or degrade performance in ML portfolios.
Key findings
- Missingness occurs in large time-organized blocks.
- Cross-sectional correlations among predictors are small.
- Missingness tends to occur in blocks organized by the underlying data source.
- Cross-sectional mean imputation performs well compared to EM methods in this context.
- Sophisticated imputations can introduce estimation noise that harms performance if ML methods are not carefully applied.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.