Skip to main content
QUICK REVIEW

[論文レビュー] Comparing hundreds of machine learning classifiers and discrete choice models in predicting travel behavior: an empirical benchmark

Shenhao Wang, Baichuan Mo|arXiv (Cornell University)|Feb 1, 2021
Transportation Planning and OptimizationSocial Sciences参考文献 72被引用数 21
ひとこと要約

本研究は、4つのハイパードメイン(モデルファミリー、データセット、サンプルサイズ、出力)をカバーする6,970回の実験を通じて、105種類の機械学習(ML)および離散選択モデル(DCM)分類器を評価することにより、これまでで最も包括的な実証的ベンチマークを提供している。その結果、アンサンブル手法(例:ランダムフォレスト、ブースティング)および深層ニューラルネットワークが予測精度が最も高く、ランダムフォレストは性能と計算効率のバランスが最も優れていることが判明した。一方、DCMはわずかに精度が低いものの、スケールが大きくなると計算的に非現実的であることが明らかになった。

ABSTRACT

Numerous studies have compared machine learning (ML) and discrete choice models (DCMs) in predicting travel demand. However, these studies often lack generalizability as they compare models deterministically without considering contextual variations. To address this limitation, our study develops an empirical benchmark by designing a tournament model, thus efficiently summarizing a large number of experiments, quantifying the randomness in model comparisons, and using formal statistical tests to differentiate between the model and contextual effects. This benchmark study compares two large-scale data sources: a database compiled from literature review summarizing 136 experiments from 35 studies, and our own experiment data, encompassing a total of 6,970 experiments from 105 models and 12 model families. This benchmark study yields two key findings. Firstly, many ML models, particularly the ensemble methods and deep learning, statistically outperform the DCM family (i.e., multinomial, nested, and mixed logit models). However, this study also highlights the crucial role of the contextual factors (i.e., data sources, inputs and choice categories), which can explain models' predictive performance more effectively than the differences in model types alone. Model performance varies significantly with data sources, improving with larger sample sizes and lower dimensional alternative sets. After controlling all the model and contextual factors, significant randomness still remains, implying inherent uncertainty in such model comparisons. Overall, we suggest that future researchers shift more focus from context-specific model comparisons towards examining model transferability across contexts and characterizing the inherent uncertainty in ML, thus creating more robust and generalizable next-generation travel demand models.

研究の動機と目的

  • 移動行動予測におけるMLおよびDCM分類器を比較する、決定的で一般化可能な実証的ベンチマークを確立すること。
  • モデルの性能がデータセット、サンプルサイズ、出力タイプによってどのように変化するかを調査すること。
  • 多様なモデルファミリーにおける予測精度と計算コストのトレードオフを評価すること。
  • 今後の研究を支援するため、高性能なモデルを同定し、DCMの計算効率の向上を提案すること。
  • 共有の公開データセットと標準化されたベンチマークフレームワークの導入を提唱することで、研究手法の一貫性を促進すること。

提案手法

  • 本研究は、12のファミリーに属する105種類の分類器、3つのデータセット(NHTS2017、LTDS2015、および追加の1つのデータセット)、3つのサンプルサイズ、3つの出力タイプ(二値、多項、順序選択)をカバーする広範な実験空間を構築した。
  • 各実験ポイントは、固定されたハイパードメインを持つ訓練済みモデルを表し、合計6,970件の固有の実験が得られた。
  • 予測精度は標準指標(例:分類精度)を用いて測定され、計算コストは学習時間を記録することで評価された。
  • 妥当性と一般化性を確認するため、35件の先行研究から得られた136件の実験ポイントを含むメタデータセットを用いて結果を検証した。
  • このフレームワークにより、将来的に新たな研究が新たな実験ポイントとして追加可能となり、継続的なベンチマークと知識蓄積が可能になる。

実験結果

リサーチクエスチョン

  • RQ1移動行動モデリングにおいて、どの機械学習および離散選択モデルファミリーが予測精度が最も高いか?
  • RQ2モデルの性能は、異なるデータセット、サンプルサイズ、出力タイプによってどのように変化するか?
  • RQ3MLおよびDCM分類器において、予測精度と計算コストのトレードオフはどのようなものか?
  • RQ4異なる実験条件下で分類器の相対的順位はどの程度安定しているか?
  • RQ5離散選択モデルの計算効率をどの程度向上させれば、ビッグデータ応用に実用的になるか?

主な発見

  • アンサンブル手法(ランダムフォレスト、勾配ブースティング、バギングを含む)は、評価されたすべての分類器の中で予測精度が最も高かった。
  • 深層ニューラルネットワーク(DNN)も上位のパフォーマンスを達成したが、著しく高い計算リソースを要した。
  • ランダムフォレストは、予測精度と計算効率の両面で最良のバランスを示し、理想的なベースラインモデルである。
  • 離散選択モデル(DCM)は、トップクラスのMLモデルと比較して3〜4パーセンテージポイント低い精度であったが、特に大規模なデータセットや高次元入力では計算が著しく遅いことが判明した。
  • 分類器の相対的順位は、データセットや条件に関わらず極めて安定しており、一方で絶対的な精度や計算時間は大きく変動した。
  • DCMはビッグデータ環境下で深刻な計算的ボトル neck を抱えており、DCMコミュニティがモデル適合よりも計算効率の向上を最優先にすべきであると示唆された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。