[论文解读] Comparing hundreds of machine learning classifiers and discrete choice models in predicting travel behavior: an empirical benchmark
本研究通过在四个超维度(模型族、数据集、样本量和输出类型)上对105种机器学习(ML)和离散选择模型(DCM)分类器进行6,970次实验,提供了迄今为止最全面的实证基准。研究发现,集成方法(如随机森林、提升法)和深度神经网络在预测准确性方面表现最佳,其中随机森林在性能与计算效率之间实现了最佳平衡;而DCM虽然略逊色于前者,但在大规模应用中仍面临计算瓶颈。
Numerous studies have compared machine learning (ML) and discrete choice models (DCMs) in predicting travel demand. However, these studies often lack generalizability as they compare models deterministically without considering contextual variations. To address this limitation, our study develops an empirical benchmark by designing a tournament model, thus efficiently summarizing a large number of experiments, quantifying the randomness in model comparisons, and using formal statistical tests to differentiate between the model and contextual effects. This benchmark study compares two large-scale data sources: a database compiled from literature review summarizing 136 experiments from 35 studies, and our own experiment data, encompassing a total of 6,970 experiments from 105 models and 12 model families. This benchmark study yields two key findings. Firstly, many ML models, particularly the ensemble methods and deep learning, statistically outperform the DCM family (i.e., multinomial, nested, and mixed logit models). However, this study also highlights the crucial role of the contextual factors (i.e., data sources, inputs and choice categories), which can explain models' predictive performance more effectively than the differences in model types alone. Model performance varies significantly with data sources, improving with larger sample sizes and lower dimensional alternative sets. After controlling all the model and contextual factors, significant randomness still remains, implying inherent uncertainty in such model comparisons. Overall, we suggest that future researchers shift more focus from context-specific model comparisons towards examining model transferability across contexts and characterizing the inherent uncertainty in ML, thus creating more robust and generalizable next-generation travel demand models.
研究动机与目标
- 建立一个确定的、可推广的实证基准,用于比较出行行为预测中的ML和DCM分类器。
- 探究模型性能在不同数据集、样本量和输出类型下的变化规律。
- 评估不同模型族在预测准确性与计算成本之间的权衡关系。
- 通过识别高性能模型并为DCM提出计算效率改进建议,为未来研究提供指导。
- 通过倡导共享公共数据集和标准化基准框架,推动方法论的一致性。
提出的方法
- 本研究构建了一个涵盖12个模型族、3个数据集(NHTS2017、LTDS2015及一个附加数据集)、3个样本量和3种输出类型(二值、多项式及有序选择)的广泛实验空间,共包含105种分类器。
- 每个实验点代表一个具有固定超维度的训练模型,共产生6,970个独立实验。
- 预测准确性采用标准指标(如分类准确率)进行衡量,计算成本以训练时间记录。
- 通过来自35项先前研究的136个实验点的元数据集对结果进行验证,以确保稳健性和可推广性。
- 该框架支持未来研究以新实验点形式持续加入,实现持续基准测试与知识积累。
实验结果
研究问题
- RQ1在出行行为建模中,哪些机器学习和离散选择模型族能够实现最高的预测准确性?
- RQ2模型性能在不同数据集、样本量和输出类型之间如何变化?
- RQ3ML和DCM分类器在预测准确性与计算成本之间存在何种权衡?
- RQ4在不同实验条件下,分类器的相对排名是否稳定?
- RQ5在计算效率方面进行哪些改进,可使离散选择模型适用于大数据应用?
主要发现
- 集成方法(包括随机森林、梯度提升和装袋法)在所有评估的分类器中实现了最高的预测准确性。
- 深度神经网络(DNNs)也表现出顶级性能,但需要显著更高的计算资源。
- 随机森林在预测准确性和计算效率之间实现了最佳平衡,是理想的基线模型。
- 离散选择模型(DCMs)比顶级ML模型低3–4个百分点的准确率,且计算速度慢得多,尤其在处理大规模数据集或高维输入时更为明显。
- 尽管绝对准确率和计算时间存在显著差异,分类器的相对排名在不同数据集和实验条件下保持高度稳定。
- DCMs在大数据场景下面临关键的计算瓶颈,提示DCM研究社区应优先关注计算效率,而非模型拟合优化。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。