[论文解读] Federated Learning on the Road: Autonomous Controller Design for Connected and Autonomous Vehicles
该论文提出了一种用于训练自动驾驶车辆控制器的动态联邦近端(DFP)学习框架,通过联网自动驾驶车辆(CAVs)之间的协作式、去中心化学习实现。考虑到车辆移动性、无线信道衰落以及非独立同分布(non-iid)、不平衡的数据,DFP算法相比FedAvg和FedProx实现了40%更快的收敛速度,其有效性基于真实车载数据轨迹得到验证。
A new federated learning (FL) framework enabled by large-scale wireless connectivity is proposed for designing the autonomous controller of connected and autonomous vehicles (CAVs). In this framework, the learning models used by the controllers are collaboratively trained among a group of CAVs. To capture the varying CAV participation in the FL training process and the diverse local data quality among CAVs, a novel dynamic federated proximal (DFP) algorithm is proposed that accounts for the mobility of CAVs, the wireless fading channels, as well as the unbalanced and nonindependent and identically distributed data across CAVs. A rigorous convergence analysis is performed for the proposed algorithm to identify how fast the CAVs converge to using the optimal autonomous controller. In particular, the impacts of varying CAV participation in the FL process and diverse CAV data quality on the convergence of the proposed DFP algorithm are explicitly analyzed. Leveraging this analysis, an incentive mechanism based on contract theory is designed to improve the FL convergence speed. Simulation results using real vehicular data traces show that the proposed DFP-based controller can accurately track the target CAV speed over time and under different traffic scenarios. Moreover, the results show that the proposed DFP algorithm has a much faster convergence compared to popular FL algorithms such as federated averaging (FedAvg) and federated proximal (FedProx). The results also validate the feasibility of the contract-theoretic incentive mechanism and show that the proposed mechanism can improve the convergence speed of the DFP algorithm by 40% compared to the baselines.
研究动机与目标
- 解决设计鲁棒的自动驾驶控制器以适应多样化、动态交通与道路条件的挑战。
- 克服传统反馈控制与基于本地学习的控制器在数据稀缺和环境变化下失效的局限性。
- 通过联邦学习实现CAVs之间车辆控制模型的协作式、去中心化训练,以提升泛化能力与性能。
- 设计一种动态、自适应的学习框架,以应对CAV参与度变化、无线信道条件波动以及异构数据质量的挑战。
- 基于契约理论设计激励机制,以提升联邦训练过程中的收敛速度与参与度。
提出的方法
- 提出一种动态联邦近端(DFP)算法,将CAV移动性的时序动态特性与无线信道衰落特性整合进联邦学习(FL)优化过程。
- 将联邦学习训练过程建模为具有时变客户端参与度和CAVs间非独立同分布(non-i.i.i.d.)数据分布的随机优化问题。
- 在损失函数中引入近端项,以在数据异构性条件下稳定训练并减少模型偏差。
- 设计基于契约理论的激励机制,根据数据质量与贡献度,使CAVs的参与动机与全局模型收敛目标保持一致。
- 使用来自BDD100K等数据集的真实车载数据轨迹,模拟多样化的交通场景并验证控制器性能。
- 进行严格的收敛性分析,量化CAV参与度波动与数据质量对模型收敛速度的影响。
实验结果
研究问题
- RQ1联邦训练中CAV参与度的变化如何影响自动驾驶控制器学习的收敛速度?
- RQ2异构的、非独立同分布的以及低质量的本地数据对CAVs中联邦学习性能有何影响?
- RQ3动态联邦近端算法能否在车辆网络中有效应对移动性与无线信道变化带来的训练不稳定性?
- RQ4基于契约理论的激励机制在多大程度上可提升基于联邦学习的控制器训练过程中的收敛速度与参与度?
- RQ5在多样化交通场景下,所提出的基于DFP的控制器与FedAvg和FedProx相比,在跟踪目标车速方面表现如何?
主要发现
- 在使用真实车载数据轨迹的仿真中,所提出的DFP算法相比FedAvg和FedProx实现了40%更快的收敛速度。
- DFP算法在非独立同分布数据与CAV参与度变化条件下有效稳定了训练过程,并具备理论收敛保证。
- 基于契约理论的激励机制通过协调CAVs的参与动机与全局模型性能,成功将收敛速度提升了40%。
- 基于DFP的控制器在包括启停交通与高速公路汇入等多样化交通条件下,能够准确跟踪目标车速。
- 收敛速率分析证实,数据质量与参与频率是影响联邦学习性能的关键因素,而DFP算法能有效缓解这些问题。
- 仿真结果表明,该协作学习框架在控制精度与鲁棒性方面显著优于仅本地训练的方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。