Skip to main content
QUICK REVIEW

[論文レビュー] Evaluating the Determinants of Mode Choice Using Statistical and Machine Learning Techniques in the Indian Megacity of Bengaluru

Tanmay Ghosh, Nithin Nagaraj|arXiv (Cornell University)|Jan 25, 2024
Transportation Planning and Optimization被引用数 4
ひとこと要約

本研究では、1,350世帯のデータセットを用いてバンガロールにおけるモード選択行動を統計的および機械学習モデルで評価し、多項ロジットとランダムフォレスト、XGBoost、SVMを比較した。ランダムフォレストモデルが最高の精度(テストデータで60.5%)を達成し、解釈可能性技術により、移動コストが10%上昇するとバス利用の確率が0.34–0.66%低下するのに対し、時間の10%短縮はマーティー利用の確率を0.16–0.42%上昇させることが明らかになった。

ABSTRACT

The decision making involved behind the mode choice is critical for transportation planning. While statistical learning techniques like discrete choice models have been used traditionally, machine learning (ML) models have gained traction recently among the transportation planners due to their higher predictive performance. However, the black box nature of ML models pose significant interpretability challenges, limiting their practical application in decision and policy making. This study utilised a dataset of $1350$ households belonging to low and low-middle income bracket in the city of Bengaluru to investigate mode choice decision making behaviour using Multinomial logit model and ML classifiers like decision trees, random forests, extreme gradient boosting and support vector machines. In terms of accuracy, random forest model performed the best ($0.788$ on training data and $0.605$ on testing data) compared to all the other models. This research has adopted modern interpretability techniques like feature importance and individual conditional expectation plots to explain the decision making behaviour using ML models. A higher travel costs significantly reduce the predicted probability of bus usage compared to other modes (a $0.66\%$ and $0.34\%$ reduction using Random Forests and XGBoost model for $10\%$ increase in travel cost). However, reducing travel time by $10\%$ increases the preference for the metro ($0.16\%$ in Random Forests and 0.42% in XGBoost). This research augments the ongoing research on mode choice analysis using machine learning techniques, which would help in improving the understanding of the performance of these models with real-world data in terms of both accuracy and interpretability.

研究の動機と目的

  • バンガロールの低所得および低中所得世帯におけるモード選択の決定要因を分析すること。
  • 従来の統計モデル(例:多項ロジット)と現代の機械学習分類器の予測性能を比較すること。
  • 特徴量の重要度と個別条件期待値(ICE)プロットを用いた機械学習モデルの解釈可能性を評価すること。
  • 解釈可能な機械学習技術を用いて、移動コストおよび時間の変化がモード選好に与える影響を定量化すること。
  • 実世界の都市データにおける高い精度とモデルの透明性を統合することで、根拠に基づく交通政策を支援すること。

提案手法

  • バンガロールの低所得および低中所得世帯から1,350世帯のデータセットを収集した。
  • モード選択分析のベンチマークとして多項ロジットモデルを適用した。
  • 意思決定木、ランダムフォレスト、XGBoost、サポートベクターマシンの4つの機械学習分類器を訓練および比較した。
  • 特徴量の重要度と個別条件期待値(ICE)プロットを用いて、モデルの意思決定と変数の影響を解釈した。
  • 訓練データおよびテストデータのスプリットにおける正答率を用いて、モデルの性能を評価した。
  • 訓練済みの機械学習モデルを用いて、移動コストおよび時間の変化がモード選択確率に与える限界効果を定量化した。

実験結果

リサーチクエスチョン

  • RQ1統計的モデルと機械学習モデルのうち、どのモデルがバンガロールの都市的文脈におけるモード選択予測に最も優れているか?
  • RQ2移動コストおよび時間の変化が、特定の輸送モード選択の予測確率にどのように影響するか?
  • RQ3特徴量の重要度やICEプロットといった解釈可能性技術は、交通計画における機械学習モデルの透明性をどの程度向上できるか?
  • RQ4所得水準および移動時間の変数が、モード選択意思決定に与える相対的な影響は何か?
  • RQ5精度と解釈可能性の観点から、機械学習モデルの予測は、従来の離散的選択モデルと比べてどのように異なるか?

主な発見

  • ランダムフォレストモデルが、テストデータで60.5%の最高精度を達成し、多項ロジット、意思決定木、XGBoost、SVMを上回った。
  • 移動コストが10%上昇すると、バス利用の予測確率はランダムフォレストで0.66%低下し、XGBoostでは0.34%低下した。
  • 移動時間が10%短縮されると、マーティー利用の予測確率はランダムフォレストで0.16%上昇し、XGBoostでは0.42%上昇した。
  • 特徴量の重要度とICEプロットにより、移動コストと移動時間が、すべての機械学習モデルにおいて最も影響力のある予測変数であることが明らかになった。
  • 本研究では、ランダムフォレストやXGBoostのような高性能な機械学習モデルが、現代の説明可能技術を用いることで解釈可能にできることが示された。
  • 解釈可能性ツールの統合により、政策立案者が機械学習ベースの予測を理解し、信頼することが可能となり、都市交通計画に活用できるようになった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。