[论文解读] Deep Neural Networks for Choice Analysis: A Statistical Learning Theory Perspective
本文提出了一种统计学习理论框架,用于解决深度神经网络(DNNs)在选择分析中的过拟合与可解释性问题。通过分解估计误差与近似误差,并将解释损失定义为真实与估计选择概率函数之间的差异,研究证明DNNs在样本量较大(>10⁴)时,无论在预测性能还是可解释性方面,均优于二值逻辑回归模型(BNL),为传统选择模型提供了一种理论基础坚实且可解释的替代方案。
While researchers increasingly use deep neural networks (DNN) to analyze individual choices, overfitting and interpretability issues remain as obstacles in theory and practice. By using statistical learning theory, this study presents a framework to examine the tradeoff between estimation and approximation errors, and between prediction and interpretation losses. It operationalizes the DNN interpretability in the choice analysis by formulating the metrics of interpretation loss as the difference between true and estimated choice probability functions. This study also uses the statistical learning theory to upper bound the estimation error of both prediction and interpretation losses in DNN, shedding light on why DNN does not have the overfitting issue. Three scenarios are then simulated to compare DNN to binary logit model (BNL). We found that DNN outperforms BNL in terms of both prediction and interpretation for most of the scenarios, and larger sample size unleashes the predictive power of DNN but not BNL. DNN is also used to analyze the choice of trip purposes and travel modes based on the National Household Travel Survey 2017 (NHTS2017) dataset. These experiments indicate that DNN can be used for choice analysis beyond the current practice of demand forecasting because it has the inherent utility interpretation, the flexibility of accommodating various information formats, and the power of automatically learning utility specification. DNN is both more predictive and interpretable than BNL unless the modelers have complete knowledge about the choice task, and the sample size is small. Overall, statistical learning theory can be a foundation for future studies in the non-asymptotic data regime or using high-dimensional statistical models in choice analysis, and the experiments show the feasibility and effectiveness of DNN for its wide applications to policy and behavioral analysis.
研究动机与目标
- 解决在小样本、非渐近条件下DNNs用于选择分析时的过拟合问题。
- 通过函数估计与选择概率损失,形式化并实现基于DNN的选择模型的可解释性。
- 在模拟与现实场景中,比较DNN与二值逻辑回归模型(BNL)在预测与可解释性方面的表现。
- 利用统计学习理论为高维、非渐近选择建模建立理论基础。
- 证明DNNs可实现与传统模型相当或更优的行为与政策相关可解释性。
提出的方法
- 使用统计学习理论,通过Radechaker复杂度对DNN中的估计误差进行上界估计,避免依赖VC维。
- 将误差分解为估计与近似两部分,表明参数范数而非参数数量决定了估计误差的上界。
- 将解释损失操作化为真实与估计选择概率函数之间的L2距离。
- 通过三种情景的蒙特卡洛模拟,比较DNN与BNL在预测与可解释性表现上的差异。
- 将DNN应用于2017年国家家庭出行调查(NHTS2017),分析出行目的与出行方式选择,验证其在真实世界中的适用性。
- 将可解释性重新定义为以预测为导向的函数估计,而非参数估计,从而实现自动效用结构设定。
实验结果
研究问题
- RQ1在小样本、非渐近条件下,DNNs是否能可靠地用于选择建模而不会过拟合?
- RQ2如何正式定义并度量基于DNN的选择模型中的可解释性,尤其是与传统选择模型相比?
- RQ3在不同数据条件下,DNNs在预测准确率与可解释性方面相对于BNL的优越程度如何?
- RQ4样本量在多大程度上影响DNNs在选择分析中相对于传统模型的预测能力释放?
- RQ5DNNs是否能自动学习到与手工设计相当或更优的行为相关效用结构?
主要发现
- 在大多数模拟情景中,DNNs在预测与可解释性方面均优于二值逻辑回归模型(BNL),尤其当样本量超过10⁴时表现更优。
- DNNs的预测优势仅在样本量较大时显现,而BNL的性能无论样本量大小均趋于饱和。
- 通过Radechaker复杂度有效界定了解释损失,表明当控制模型范数时,DNNs在非渐近条件下不会过拟合。
- DNNs通过估计完整的选择概率函数,而非依赖单个参数估计,实现了与BNL相当或更优的可解释性。
- 在NHTS2017的应用中,DNNs成功建模了出行目的与出行方式选择,具备高预测能力与可解释的效用结构。
- 研究表明,DNNs可自动学习复杂的效用结构,从而减少对基于领域知识的特征工程的依赖。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。