Skip to main content
QUICK REVIEW

[论文解读] Machine Learning for Economics Research: When What and How?

Ajit Desai|arXiv (Cornell University)|Mar 31, 2023
Stock Market Forecasting MethodsDecision Sciences被引用 3
一句话总结

本文对经济学中机器学习(ML)应用进行了精心梳理的综述,探讨了ML的应用时机、优选模型及其应用方式。研究发现,ML在处理非传统数据、捕捉非线性关系以及提升预测性能方面表现优异,尤其在处理非结构化数据时深度学习表现突出,而在传统数据集上则集成方法更为有效;同时强调应根据需求定制模型并采用迁移学习以获得最佳效果。

ABSTRACT

This article provides a curated review of selected papers published in prominent economics journals that use machine learning (ML) tools for research and policy analysis. The review focuses on three key questions: (1) when ML is used in economics, (2) what ML models are commonly preferred, and (3) how they are used for economic applications. The review highlights that ML is particularly used to process nontraditional and unstructured data, capture strong nonlinearity, and improve prediction accuracy. Deep learning models are suitable for nontraditional data, whereas ensemble learning models are preferred for traditional datasets. While traditional econometric models may suffice for analyzing low-complexity data, the increasing complexity of economic data due to rapid digitalization and the growing literature suggests that ML is becoming an essential addition to the econometrician's toolbox.

研究动机与目标

  • 为经济学家和数据科学家提供指导,帮助其在经济研究与政策分析中有效应用机器学习工具。
  • 明确在数字化带来数据复杂性日益增加的背景下,机器学习与传统计量经济学方法相比在何种情境下更为适用。
  • 基于顶级经济学期刊中的实证应用,识别适用于不同数据类型和研究需求的最有效ML模型。
  • 突出模型选择、定制化以及迁移学习的最佳实践,以提升模型性能与可解释性。
  • 探讨机器学习在经济学中面临的关键局限,如可解释性不足、数据需求高以及缺乏标准误,并指出新兴解决方案。

提出的方法

  • 从10本领先的经济学期刊(如AER、QJE、JPE)中精选论文,通过关键词搜索(如'machine learning'、'ensemble learning'、'deep learning'、'reinforcement learning'、'NLP')进行系统梳理。
  • 根据数据类型(传统数据与非传统数据)、模型类型(深度学习、集成模型、因果ML)和应用场景(预测、特征提取、因果推断)对ML应用进行分类。
  • 分析模型选择模式:文本/音频/图像数据偏好使用深度学习模型;表格数据则更适用集成模型(如随机森林、XGBoost);在非传统数据有限时,迁移学习成为关键策略。
  • 利用文章标题与摘要中的词云图,可视化ML应用中的关键词与趋势(如'data'、'effect'、'decision'、'machine learning')。
  • 结合近期文献中的案例研究,说明实际应用:自编码器在资产定价中的应用、谷歌关键词拍卖的无监督聚类、中央银行中的强化学习。
  • 评估可解释性技术,如基于Shapley值的方法(如SHAP),以及正则化ML模型的渐近理论发展。
Figure 1: The number of publications over five years (between 2018-2022) in the leading economics journals that use ML. The data includes articles from the following ten journals: American Economic Review (AER), Econometrica, Journal of Economic Perspectives (JEP), Journal of Monetary Economics (JME
Figure 1: The number of publications over five years (between 2018-2022) in the leading economics journals that use ML. The data includes articles from the following ten journals: American Economic Review (AER), Econometrica, Journal of Economic Perspectives (JEP), Journal of Monetary Economics (JME

实验结果

研究问题

  • RQ1在经济学研究中,机器学习在何种情况下最有益,特别是相较于传统计量经济学模型?
  • RQ2经济学中常用的机器学习模型有哪些?模型选择如何随数据特征而变化?
  • RQ3如何将机器学习有效应用于文本、图像和音频等非传统数据,以支持经济分析?
  • RQ4将ML应用于经济学面临的主要挑战是什么?研究人员如何应对可解释性、偏差以及缺乏标准误等问题?
  • RQ5当数据有限但复杂时,迁移学习和预训练模型能在多大程度上提升性能?

主要发现

  • 机器学习在经济学中日益被用于处理非传统和非结构化数据,捕捉强非线性关系,并在预测精度上超越传统计量经济学模型。
  • 对于文本、音频和图像数据,Transformer和ConvNext等深度学习模型更受青睐;而对于传统表格数据,集成学习模型(如随机森林、XGBoost)表现最为有效。
  • 使用预训练模型的迁移学习显著提升了在非传统数据有限情况下的性能,是数据稀缺应用中的关键策略。
  • 因果机器学习模型正在兴起,用于因果推断,但目前仍不如预测模型普遍。
  • 可解释性仍是主要挑战;基于Shapley值的方法(如SHAP)被用于解释模型输出,但尚缺乏正式的渐近理论支持。
  • 尽管已有进展,ML模型仍缺乏标准误和渐近性质,过拟合与数据偏差在经济学应用中仍是关键关切。
Figure 2: Schematic diagram representing the relative merits of ML and traditional econometric methods. The plot is adapted from [ 1 , 18 ] .
Figure 2: Schematic diagram representing the relative merits of ML and traditional econometric methods. The plot is adapted from [ 1 , 18 ] .

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。