Skip to main content
QUICK REVIEW

[论文解读] Choice modelling in the age of machine learning

Sander van Cranenburgh, S. Wang|arXiv (Cornell University)|Jan 1, 2021
Forecasting Techniques and Applications参考文献 112被引用 9
一句话总结

本文通过对比基于理论的范式与数据驱动的范式,识别出二者之间的协同效应,倡导将机器学习更深层次地融入选择建模,提出一种混合未来:机器学习可增强模型的灵活性和数据处理能力,尤其在处理文本和图像输入方面,同时保留传统模型的可解释性和理论基础。

ABSTRACT

Since its inception, the choice modelling field has been dominated by theory-driven models. The recent emergence and growing popularity of machine learning models offer an alternative data-driven approach. Machine learning models, techniques and practices could help overcome problems and limitations of the current theory-driven modelling paradigm, e.g. relating to the ad-hocness in search for the optimal model specification, and theory-driven choice model's inability to work with text and image data. However, despite the potential value of machine learning to improve choice modelling practices, the choice modelling field has been somewhat hesitant to embrace machine learning. The aim of this paper is to facilitate (further) integration of machine learning in the choice modelling field. To achieve this objective, we make the case that (further) integration of machine learning in the choice modelling field is beneficial for the choice modelling field, and, we shed light on where the benefits of further integration can be found. Specifically, we take the following approach. First, we clarify the similarities and differences between the two modelling paradigms. Second, we provide a literature overview on the use of machine learning for choice modelling. Third, we reinforce the strengths of the current theory-driven modelling paradigm and compare this with the machine learning modelling paradigm, Fourth, we identify opportunities for embracing machine learning for choice modelling, while recognising the strengths of the current theory-driven paradigm. Finally, we put forward a vision on the future relationship between the theory-driven choice models and machine learning.

研究动机与目标

  • 解决该领域尽管具备潜力,却对采用机器学习持犹豫态度的问题,以克服基于理论的选择模型的局限性。
  • 厘清基于理论的选择模型与机器学习方法之间的差异与协同效应。
  • 识别机器学习可增强选择建模的具体机会,特别是在处理文本和图像等复杂数据类型方面。
  • 在提出一种平衡且整合的未来方向的同时,强化基于理论的模型的持久价值。

提出的方法

  • 对基于理论的建模范式与机器学习建模范式进行对比分析,重点关注假设、可解释性及数据需求。
  • 对现有文献中机器学习在不同领域选择建模中的应用进行系统性综述。
  • 通过结构化对比,识别出两种范式的关键优势,例如传统模型的理论一致性与机器学习在数据灵活性方面的优势。
  • 提出一种混合建模愿景,即利用机器学习技术增强而非取代传统选择模型。
  • 运用概念框架,映射机器学习在模型设定、数据处理和预测性能方面可提升的领域。
  • 强调方法论的整合,包括利用机器学习进行特征工程和选择数据中非线性模式的检测。

实验结果

研究问题

  • RQ1基于理论的选择模型与机器学习模型在假设和结构上存在哪些差异?
  • RQ2机器学习技术在何种方式下可改善选择模型的设定与性能,特别是在处理文本和图像等非传统数据时?
  • RQ3当前基于理论的模型存在哪些关键局限,而机器学习可予以解决?
  • RQ4在整合数据驱动的机器学习方法时,如何保持传统模型的可解释性和理论基础?
  • RQ5未来何种模型架构能够最优地结合两种范式的优点于选择建模之中?

主要发现

  • 机器学习模型在处理复杂、高维数据(如文本和图像)方面具有显著优势,而传统基于理论的模型则难以应对。
  • 当前基于理论的范式常因模型设定过程的随意性而受限,而机器学习可通过自动化特征学习与超参数调优来缓解此问题。
  • 尽管具有数据驱动特性,机器学习模型仍可通过SHAP或LIME等技术提高可解释性,从而在一定程度上满足选择建模对可解释性的需求。
  • 将机器学习整合到选择建模中并非取代基于理论的模型,而是作为互补增强,尤其在模型发现与数据预处理方面。
  • 混合未来模型是可行且有益的:机器学习负责处理数据复杂性与非线性关系,而基于理论的模型则确保可解释性与因果推断能力。
  • 关于机器学习在选择建模中应用的文献正在增长,表明其潜力正日益获得认可,但主流选择建模研究中仍存在全面整合的不足。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。