Skip to main content
QUICK REVIEW

[论文解读] Evaluating the Determinants of Mode Choice Using Statistical and Machine Learning Techniques in the Indian Megacity of Bengaluru

Tanmay Ghosh, Nithin Nagaraj|arXiv (Cornell University)|Jan 25, 2024
Transportation Planning and Optimization被引用 4
一句话总结

本研究基于1,350户家庭的数据集,采用统计模型与机器学习模型评估班加鲁鲁的出行方式选择行为,对比了多项Logit模型与随机森林、XGBoost及SVM模型。随机森林模型在测试数据上的准确率最高(60.5%),可解释性技术分析显示,出行成本每增加10%,公交车使用概率下降0.34%–0.66%;出行时间每减少10%,地铁偏好概率上升0.16%–0.42%。

ABSTRACT

The decision making involved behind the mode choice is critical for transportation planning. While statistical learning techniques like discrete choice models have been used traditionally, machine learning (ML) models have gained traction recently among the transportation planners due to their higher predictive performance. However, the black box nature of ML models pose significant interpretability challenges, limiting their practical application in decision and policy making. This study utilised a dataset of $1350$ households belonging to low and low-middle income bracket in the city of Bengaluru to investigate mode choice decision making behaviour using Multinomial logit model and ML classifiers like decision trees, random forests, extreme gradient boosting and support vector machines. In terms of accuracy, random forest model performed the best ($0.788$ on training data and $0.605$ on testing data) compared to all the other models. This research has adopted modern interpretability techniques like feature importance and individual conditional expectation plots to explain the decision making behaviour using ML models. A higher travel costs significantly reduce the predicted probability of bus usage compared to other modes (a $0.66\%$ and $0.34\%$ reduction using Random Forests and XGBoost model for $10\%$ increase in travel cost). However, reducing travel time by $10\%$ increases the preference for the metro ($0.16\%$ in Random Forests and 0.42% in XGBoost). This research augments the ongoing research on mode choice analysis using machine learning techniques, which would help in improving the understanding of the performance of these models with real-world data in terms of both accuracy and interpretability.

研究动机与目标

  • 分析班加鲁鲁低收入及低中等收入家庭出行方式选择的影响因素。
  • 比较传统统计模型(如多项Logit)与现代机器学习分类器的预测性能。
  • 利用特征重要性与个体条件期望(ICE)图,评估机器学习模型的可解释性。
  • 运用可解释性机器学习技术量化出行成本与时间变化对出行方式偏好的影响。
  • 通过结合高准确率与模型透明度,为基于证据的交通政策提供支持,应用于真实城市数据。

提出的方法

  • 从班加鲁鲁低收入及低中等收入群体中收集了1,350户家庭的数据集。
  • 采用多项Logit模型作为出行方式分析的基准统计方法。
  • 训练并比较了四种机器学习分类器:决策树、随机森林、XGBoost与支持向量机(SVM)。
  • 使用特征重要性与个体条件期望(ICE)图解释模型决策,评估变量影响。
  • 通过训练集与测试集的准确率评估模型性能。
  • 利用训练好的机器学习模型量化出行成本与时间变化对出行方式选择概率的边际效应。

实验结果

研究问题

  • RQ1在班加鲁鲁的城市背景下,统计模型与机器学习模型中哪一种在预测出行方式选择方面表现最佳?
  • RQ2出行成本与时间的变化如何影响特定交通方式选择的预测概率?
  • RQ3特征重要性与ICE图等可解释性技术在多大程度上能提升机器学习模型在交通规划中的透明度?
  • RQ4社会经济因素与出行时间变量对出行方式选择决策的相对影响是什么?
  • RQ5与传统离散选择模型相比,机器学习模型在准确率与可解释性方面有何差异?

主要发现

  • 随机森林模型在测试集上取得了最高的准确率(60.5%),优于多项Logit、决策树、XGBoost与SVM模型。
  • 出行成本增加10%,公交车使用概率的预测值下降0.66%(随机森林)与0.34%(XGBoost)。
  • 出行时间减少10%,选择地铁的预测概率上升0.16%(随机森林)与0.42%(XGBoost)。
  • 特征重要性与ICE图显示,出行成本与时间在所有机器学习模型中均为最具影响力的预测变量。
  • 本研究证明,通过现代可解释性技术,高性能机器学习模型(如随机森林与XGBoost)可实现可解释性。
  • 将可解释性工具整合到模型中,有助于政策制定者理解并信任基于机器学习的出行预测结果,用于城市交通规划。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。