Skip to main content
QUICK REVIEW

[Paper Review] Comparing hundreds of machine learning classifiers and discrete choice models in predicting travel behavior: an empirical benchmark

Shenhao Wang, Baichuan Mo|arXiv (Cornell University)|Feb 1, 2021
Transportation Planning and Optimization72 references21 citations
TL;DR

This study provides the most comprehensive empirical benchmark to date by evaluating 105 machine learning (ML) and discrete choice model (DCM) classifiers across 6,970 experiments spanning four hyper-dimensions: model families, datasets, sample sizes, and outputs. It finds that ensemble methods (e.g., random forests, boosting) and deep neural networks achieve the highest prediction accuracy, with random forests offering the best balance of performance and computational efficiency, while DCMs, though slightly less accurate, remain computationally prohibitive at scale.

ABSTRACT

Numerous studies have compared machine learning (ML) and discrete choice models (DCMs) in predicting travel demand. However, these studies often lack generalizability as they compare models deterministically without considering contextual variations. To address this limitation, our study develops an empirical benchmark by designing a tournament model, thus efficiently summarizing a large number of experiments, quantifying the randomness in model comparisons, and using formal statistical tests to differentiate between the model and contextual effects. This benchmark study compares two large-scale data sources: a database compiled from literature review summarizing 136 experiments from 35 studies, and our own experiment data, encompassing a total of 6,970 experiments from 105 models and 12 model families. This benchmark study yields two key findings. Firstly, many ML models, particularly the ensemble methods and deep learning, statistically outperform the DCM family (i.e., multinomial, nested, and mixed logit models). However, this study also highlights the crucial role of the contextual factors (i.e., data sources, inputs and choice categories), which can explain models' predictive performance more effectively than the differences in model types alone. Model performance varies significantly with data sources, improving with larger sample sizes and lower dimensional alternative sets. After controlling all the model and contextual factors, significant randomness still remains, implying inherent uncertainty in such model comparisons. Overall, we suggest that future researchers shift more focus from context-specific model comparisons towards examining model transferability across contexts and characterizing the inherent uncertainty in ML, thus creating more robust and generalizable next-generation travel demand models.

Motivation & Objective

  • To establish a definitive, generalizable empirical benchmark for comparing ML and DCM classifiers in travel behavior prediction.
  • To investigate how model performance varies across datasets, sample sizes, and output types.
  • To evaluate the trade-offs between prediction accuracy and computational cost across diverse model families.
  • To guide future research by identifying high-performing models and suggesting improvements in computational efficiency for DCMs.
  • To promote methodological consistency by advocating for shared public datasets and a standardized benchmark framework.

Proposed method

  • The study constructs an extensive experiment space spanning 105 classifiers from 12 families, 3 datasets (NHTS2017, LTDS2015, and one additional dataset), 3 sample sizes, and 3 output types (binary, multinomial, and ordered choice).
  • Each experiment point represents a trained model with fixed hyper-dimensions, resulting in 6,970 unique experiments.
  • Prediction accuracy is measured using standard metrics (e.g., classification accuracy), and computational cost is recorded as training time.
  • The results are validated using a meta-dataset of 136 experiment points from 35 prior studies to ensure robustness and generalizability.
  • The framework enables future studies to be added as new experiment points, allowing continuous benchmarking and knowledge accumulation.

Experimental results

Research questions

  • RQ1Which machine learning and discrete choice model families achieve the highest prediction accuracy in travel behavior modeling?
  • RQ2How does model performance vary across different datasets, sample sizes, and output types?
  • RQ3What is the trade-off between prediction accuracy and computational cost across ML and DCM classifiers?
  • RQ4How stable is the relative ranking of classifiers across different experimental conditions?
  • RQ5What improvements in computational efficiency could make discrete choice models viable for big data applications?

Key findings

  • Ensemble methods, including random forests, gradient boosting, and bagging, achieve the highest prediction accuracy among all classifiers evaluated.
  • Deep neural networks (DNNs) also achieve top-tier performance but require significantly higher computational resources.
  • Random forests deliver the best balance between prediction accuracy and computational efficiency, making them ideal baseline models.
  • Discrete choice models (DCMs) are 3–4 percentage points less accurate than top ML models but are computationally much slower, especially with large datasets or high-dimensional inputs.
  • The relative ranking of classifiers remains highly stable across datasets and conditions, though absolute accuracy and computational time vary widely.
  • DCMs face a critical computational bottleneck in big data contexts, suggesting a need for the DCM community to prioritize computational efficiency over model fitting.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.