[论文解读] Theory-based residual neural networks: A synergy of discrete choice models and deep neural networks
该论文提出理论基础残差神经网络(TB-ResNets),一种新颖的框架,通过(δ, 1−δ)加权方案结合离散选择模型(DCMs)与深度神经网络(DNNs)的效用函数,实现两者的协同作用。该方法通过利用DCMs稳定效用函数、DNNs捕捉复杂行为模式,显著提升了预测准确性、可解释性与鲁棒性,优于纯DCMs与DNNs。
Researchers often treat data-driven and theory-driven models as two disparate or even conflicting methods in travel behavior analysis. However, the two methods are highly complementary because data-driven methods are more predictive but less interpretable and robust, while theory-driven methods are more interpretable and robust but less predictive. Using their complementary nature, this study designs a theory-based residual neural network (TB-ResNet) framework, which synergizes discrete choice models (DCMs) and deep neural networks (DNNs) based on their shared utility interpretation. The TB-ResNet framework is simple, as it uses a ($δ$, 1-$δ$) weighting to take advantage of DCMs' simplicity and DNNs' richness, and to prevent underfitting from the DCMs and overfitting from the DNNs. This framework is also flexible: three instances of TB-ResNets are designed based on multinomial logit model (MNL-ResNets), prospect theory (PT-ResNets), and hyperbolic discounting (HD-ResNets), which are tested on three data sets. Compared to pure DCMs, the TB-ResNets provide greater prediction accuracy and reveal a richer set of behavioral mechanisms owing to the utility function augmented by the DNN component in the TB-ResNets. Compared to pure DNNs, the TB-ResNets can modestly improve prediction and significantly improve interpretation and robustness, because the DCM component in the TB-ResNets stabilizes the utility functions and input gradients. Overall, this study demonstrates that it is both feasible and desirable to synergize DCMs and DNNs by combining their utility specifications under a TB-ResNet framework. Although some limitations remain, this TB-ResNet framework is an important first step to create mutual benefits between DCMs and DNNs for travel behavior modeling, with joint improvement in prediction, interpretation, and robustness.
研究动机与目标
- 解决出行行为研究中数据驱动机器学习与理论驱动离散选择模型之间的张力。
- 通过协同结合两种方法,克服纯DNNs(可解释性低、鲁棒性差)与纯DCMs(预测能力弱)的局限性。
- 开发一种灵活且可泛化的框架,将特定领域的行为理论与数据驱动学习相结合,以改善建模结果。
- 证明通过基于效用的残差学习结合DCMs与DNNs,可在预测、可解释性与鲁棒性方面实现协同提升。
提出的方法
- 提出TB-ResNet框架,通过(δ, 1−δ)加权其效用函数,将理论驱动的DCM与数据驱动的DNN相结合,模仿深度网络中的残差学习机制。
- 在DNNs中使用Softmax激活函数,确保与随机效用最大化(RUM)框架的概率兼容性,从而实现效用层面的直接协同。
- 将TB-ResNet形式化为正则化机制:DCMs通过稳定DNNs防止过拟合,而DNNs通过灵活的数据驱动效用组件丰富DCMs。
- 设计三种TB-ResNet实例:基于多项对数模型的MNL-ResNet、基于前景理论的风险偏好PT-ResNet,以及基于双曲贴现的时间偏好HD-ResNet。
- 应用统计学习理论证明正则化效果,表明纯DNNs易过拟合,而纯DCMs易欠拟合。
- 在三个数据集(新加坡数据集及Tanaka, 2010的两个数据集)上进行实证测试,评估预测、可解释性与鲁棒性指标下的性能表现。
实验结果
研究问题
- RQ1能否通过共享的效用解释,有意义地协同离散选择模型与深度神经网络?
- RQ2通过残差效用框架结合DCMs与DNNs,是否能实现超越任一方法单独表现的预测准确性提升?
- RQ3在数据驱动模型中引入理论驱动组件,在多大程度上增强了可解释性与鲁棒性?
- RQ4TB-ResNet框架能否灵活适配不同行为理论(如前景理论、双曲贴现)在选择建模中的应用?
- RQ5(δ, 1−δ)加权方案如何在模型简洁性与预测丰富性之间实现权衡?
主要发现
- TB-ResNets通过在DCMs效用函数中引入DNN衍生组件,显著提升了相对于纯DCMs的预测准确性。
- 与纯DNNs相比,TB-ResNets在预测性能上略有提升,但可解释性与鲁棒性显著增强,归因于DCM组件的稳定作用。
- MNL-ResNet、PT-ResNet与HD-ResNet三种实例在预测、可解释性与鲁棒性三项评估标准上均表现更优。
- 在三个数据集上的实证结果表明,TB-ResNets在模型稳定性与行为洞察方面持续优于或匹配纯DNNs与DCMs,且表现显著提升。
- 该框架成功将多种行为理论整合进统一的深度学习架构中,揭示了比DCMs单独使用更丰富的行为机制。
- DCM组件的正则化效应降低了DNNs的过拟合风险,而DNN组件则缓解了DCMs的欠拟合问题,验证了协同作用假设。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。