[Paper Review] Machine Learning for Economic Forecasting: An Application to China's GDP Growth
The paper evaluates a wide range of machine learning models, alone and in combination with econometric methods, for forecasting China’s quarterly GDP growth and analyzes interpretability and robustness across periods including inflection points.
This paper aims to explore the application of machine learning in forecasting Chinese macroeconomic variables. Specifically, it employs various machine learning models to predict the quarterly real GDP growth of China, and analyzes the factors contributing to the performance differences among these models. Our findings indicate that the average forecast errors of machine learning models are generally lower than those of traditional econometric models or expert forecasts, particularly in periods of economic stability. However, during certain inflection points, although machine learning models still outperform traditional econometric models, expert forecasts may exhibit greater accuracy in some instances due to experts' more comprehensive understanding of the macroeconomic environment and real-time economic variables. In addition to macroeconomic forecasting, this paper employs interpretable machine learning methods to identify the key attributive variables from different machine learning models, aiming to enhance the understanding and evaluation of their contributions to macroeconomic fluctuations.
Motivation & Objective
- Assess the forecasting performance of machine learning models for China's quarterly GDP growth relative to traditional econometric models and expert forecasts.
- Evaluate combined and mixed modeling approaches to leverage strengths of both ML and econometrics.
- Apply interpretable ML methods to identify key drivers of GDP fluctuations in China.
- Examine model performance across stable periods and inflection points (crises, COVID-19).
- Provide robustness checks and temporal analysis of predictive accuracy.
Proposed method
- Catalogue models into five groups: econometric (AR, FM), machine learning (LASSO, ridge, KRR, RF, GBDT, XGBoost), combined (FM+ML), mixed (means/medians of group models), and weighted ensembles.
- Forecast using an expanding window method (EWM) with about 20 macro variables sourced mainly from Chinese statistics and international peers.
- Formulate predictions as y_{t+h}=f(Z_t)+ε_{t+h} with Z_t including lagged y and high-dimensional X; optimize via min L(y,f(Z))+R(f,ρ).
- Leverage combined models: y_t+h from ML using latent factors F_t from X_t; X_t = ΛF_t + η_t.
- Use SHAP for interpretability to assess global and local variable importance.
- Evaluate via RMSE and MAE across out-of-sample forecasts from 1996Q1 to 2023Q4; use expanding windows and robustness analyses.

Experimental results
Research questions
- RQ1Do machine learning models, alone or in combination, yield lower forecast errors than traditional econometric models and expert forecasts for China’s GDP growth?
- RQ2How do model performances differ across periods of stability versus economic inflection points (crises, COVID-19)?
- RQ3Which variables most contribute to ML-based GDP forecasts, globally and locally, according to SHAP explanations?
- RQ4Are combined and weighted ML/econometric models more robust than single-model forecasts across different time horizons?
- RQ5What is the relative forecasting performance of ML ensembles compared to expert consensus forecasts (Longrun and Yicai) during post-2014 periods?
Key findings
- ML models and ML-based hybrids generally outperform traditional econometric models in out-of-sample GDP growth forecasts.
- During stable periods, ML models and combinations yield higher accuracy than econometric and expert forecasts.
- At inflection points, ML models often predict the direction of change but may be outperformed by expert forecasts in some cases due to macroeconomic context understanding.
- Interpretable ML (SHAP) reveals global and local driver importance for GDP fluctuations across time.
- Compared to Longrun Expert Forecasts and other benchmarks, certain ML models (e.g., RF, XGBoost) achieve lower RMSE in the 2005–2015 window; combined models can perform competitively but sometimes lag expert forecasts.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.