Skip to main content
QUICK REVIEW

[论文解读] Modeling Stated Preference for Mobility-on-Demand Transit: A Comparison of Machine Learning and Logit Models

Xilei Zhao, Xiang Yan|arXiv (Cornell University)|Nov 4, 2018
Economic and Environmental Valuation参考文献 40被引用 18
一句话总结

本研究比较了机器学习(ML)模型与传统logit模型——特别是多项logit(MNL)和混合logit——在预测按需出行服务中陈述偏好(SP)时的表现。基于SP调查数据,研究结果表明,随机森林(RF)及其他ML模型在样本外预测准确性方面显著优于logit模型,同时在行为解释上也保持高度一致,且ML模型还能自动捕捉非线性关系。

ABSTRACT

Logit models are usually applied when studying individual travel behavior, i.e., to predict travel mode choice and to gain behavioral insights on traveler preferences. Recently, some studies have applied machine learning to model travel mode choice and reported higher out-of-sample predictive accuracy than traditional logit models (e.g., multinomial logit). However, little research focuses on comparing the interpretability of machine learning with logit models. In other words, how to draw behavioral insights from the high-performance "black-box" machine-learning models remains largely unsolved in the field of travel behavior modeling. This paper aims at providing a comprehensive comparison between the two approaches by examining the key similarities and differences in model development, evaluation, and behavioral interpretation between logit and machine-learning models for travel mode choice modeling. To complement the theoretical discussions, the paper also empirically evaluates the two approaches on the stated-preference survey data for a new type of transit system integrating high-frequency fixed-route services and ridesourcing. The results show that machine learning can produce significantly higher predictive accuracy than logit models. Moreover, machine learning and logit models largely agree on many aspects of behavioral interpretations. In addition, machine learning can automatically capture the nonlinear relationship between the input features and choice outcomes. The paper concludes that there is great potential in merging ideas from machine learning and conventional statistical methods to develop refined models for travel behavior research and suggests some new research directions.

研究动机与目标

  • 为弥合在出行方式选择研究中比较机器学习与logit模型可解释性与预测性能之间的空白。
  • 评估高性能的‘黑箱’ML模型是否能产生与传统logit模型相当的可靠行为洞察。
  • 探讨在陈述偏好数据背景下,ML与logit模型在数据需求、建模方法和输出解释方面的根本差异。
  • 探索将ML的预测能力与logit模型的可解释性相结合,以提升出行行为建模的潜力。
  • 评估多种ML算法与logit模型在新型按需出行服务SP调查数据集上的表现。

提出的方法

  • 通过交叉验证,实证比较七种机器学习分类器(包括随机森林、神经网络和梯度提升)与两种logit模型(MNL和混合logit)的性能。
  • 对ML模型应用可解释性技术:变量重要性、部分依赖图和敏感性分析,以提取行为洞察。
  • 使用一项拟议的按需出行服务系统中整合固定路线与按需服务的陈述偏好调查数据。
  • 通过k折交叉验证评估模型的样本外预测准确性。
  • 从logit模型中计算边际效应和弧价格弹性,以与ML推导出的行为解释进行比较。
  • 使用准确率和ROC曲线下面积等标准指标评估模型拟合度与预测能力。

实验结果

研究问题

  • RQ1机器学习模型与传统logit模型在预测按需出行服务的陈述偏好方面表现如何?
  • RQ2机器学习模型与logit模型在关键出行属性(如等待时间、换乘次数、拼车时间)的重要性及影响方向上的一致性如何?
  • RQ3机器学习模型能否自动检测并表示出行属性与出行方式选择决策之间的非线性关系?
  • RQ4ML模型与logit模型在行为解释(如边际效应、弧价格弹性)方面存在哪些关键差异?
  • RQ5机器学习模型能否作为探索性工具,用于指导更准确且基于行为的logit模型的设定?

主要发现

  • 随机森林(RF)模型在样本外预测准确性方面表现最佳,显著优于MNL和混合logit模型。
  • 大多数机器学习模型,包括RF和神经网络,在个体和聚合预测水平上均表现出优于logit模型的预测性能。
  • 混合logit模型在预测准确性上表现不如MNL模型,可能由于过拟合,这挑战了‘更复杂的logit模型总是表现更好’的假设。
  • ML模型与logit模型在关键属性(如等待时间、换乘次数、拼车时间)的重要性及影响方向上表现出高度一致性。
  • 部分依赖图显示,随机森林模型自动捕捉了出行时间与等待时间对出行方式选择的非线性影响,表明ML在揭示复杂行为模式方面具有优势。
  • ML模型的边际效应与弧价格弹性估计与logit模型趋势一致,但存在定量差异,表明ML输出可提供互补的行为洞察。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。