[Paper Review] Machine Learning for Economics Research: When What and How?
This paper provides a curated review of machine learning (ML) applications in economics, addressing when ML is used, which models are preferred, and how they are applied. It finds ML excels in handling nontraditional data, capturing nonlinearity, and improving prediction—especially with deep learning for unstructured data and ensemble methods for traditional datasets—while emphasizing tailored model use and transfer learning for optimal results.
This article provides a curated review of selected papers published in prominent economics journals that use machine learning (ML) tools for research and policy analysis. The review focuses on three key questions: (1) when ML is used in economics, (2) what ML models are commonly preferred, and (3) how they are used for economic applications. The review highlights that ML is particularly used to process nontraditional and unstructured data, capture strong nonlinearity, and improve prediction accuracy. Deep learning models are suitable for nontraditional data, whereas ensemble learning models are preferred for traditional datasets. While traditional econometric models may suffice for analyzing low-complexity data, the increasing complexity of economic data due to rapid digitalization and the growing literature suggests that ML is becoming an essential addition to the econometrician's toolbox.
Motivation & Objective
- To guide economists and data scientists in effectively applying machine learning tools to economic research and policy analysis.
- To clarify when machine learning is appropriate versus traditional econometric methods, particularly in light of increasing data complexity from digitalization.
- To identify the most effective ML models for different data types and research needs, based on empirical applications in top economics journals.
- To highlight best practices in model selection, customization, and use of transfer learning to improve performance and interpretability.
- To address key limitations of ML in economics, such as interpretability, data requirements, and lack of standard errors, while pointing to emerging solutions.
Proposed method
- Curated review of selected papers from 10 leading economics journals (e.g., AER, QJE, JPE) using keyword searches: 'machine learning', 'ensemble learning', 'deep learning', 'reinforcement learning', 'NLP'.
- Categorization of ML applications based on data type (traditional vs. nontraditional), model type (deep learning, ensemble, causal ML), and use case (prediction, feature extraction, causal inference).
- Analysis of model selection patterns: deep learning for text/audio/images; ensemble models (e.g., random forests, XGBoost) for tabular data; transfer learning for limited nontraditional data.
- Use of word clouds from article titles and abstracts to visualize key terms and trends in ML adoption (e.g., 'data', 'effect', 'decision', 'machine learning').
- Incorporation of case studies from recent literature to illustrate applications: autoencoders in asset pricing, unsupervised clustering of Google keyword auctions, reinforcement learning in central banking.
- Evaluation of interpretability techniques such as Shapley-value-based methods (e.g., SHAP) and asymptotic theory development for regularized ML models.

Experimental results
Research questions
- RQ1When is machine learning most beneficial in economics research, particularly relative to traditional econometric models?
- RQ2Which machine learning models are most commonly used in economics, and how do model choices vary with data characteristics?
- RQ3How can machine learning be effectively applied to nontraditional data such as text, images, and audio in economic analysis?
- RQ4What are the key challenges in applying ML to economics, and how are researchers addressing issues like interpretability, bias, and lack of standard errors?
- RQ5To what extent can transfer learning and pre-trained models improve performance when data is limited but complex?
Key findings
- Machine learning is increasingly used in economics to process nontraditional and unstructured data, capture strong nonlinearity, and improve prediction accuracy beyond traditional econometric models.
- Deep learning models such as Transformers and ConvNext are preferred for text, audio, and image data, while ensemble learning models (e.g., random forests, XGBoost) are most effective for traditional tabular datasets.
- Transfer learning with pre-trained models significantly enhances performance on limited nontraditional data, making it a key strategy for data-scarce applications.
- Causal machine learning models are emerging for causal inference, though they remain less common than predictive models.
- Interpretability remains a major challenge; Shapley-value-based methods (e.g., SHAP) are being used to explain model outputs, though they lack formal asymptotic theory.
- Despite progress, ML models still lack standard errors and asymptotic properties, and overfitting and data bias remain critical concerns in economic applications.
![Figure 2: Schematic diagram representing the relative merits of ML and traditional econometric methods. The plot is adapted from [ 1 , 18 ] .](https://ar5iv.labs.arxiv.org/html/2304.00086/assets/ml_vs_econometrics.png)
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.