[论文解读] 5G Network on Wings: A Deep Reinforcement Learning Approach to the UAV-based Integrated Access and Backhaul
本文提出一种去中心化的深度强化学习(DecRL)算法——DecRL-AE&VAS,利用集成接入与回传(IAB)技术,自主优化灾难救援等关键任务(MC)场景中多个无人机基站(UAV-BSs)的三维位置。通过结合自适应探索与基于价值的动作选择,该方法使UAV-BS能够动态调整位置,以应对用户移动,维持高用户吞吐量和低丢弃率,性能达到全局最优解的5%–6%以内。
Fast and reliable wireless communication has become a critical demand in human life. In the case of mission-critical (MC) scenarios, for instance, when natural disasters strike, providing ubiquitous connectivity becomes challenging by using traditional wireless networks. In this context, unmanned aerial vehicle (UAV) based aerial networks offer a promising alternative for fast, flexible, and reliable wireless communications. Due to unique characteristics such as mobility, flexible deployment, and rapid reconfiguration, drones can readily change location dynamically to provide on-demand communications to users on the ground in emergency scenarios. As a result, the usage of UAV base stations (UAV-BSs) has been considered an appropriate approach for providing rapid connection in MC scenarios. In this paper, we study how to control multiple UAV-BSs in both static and dynamic environments. We use a system-level simulator to model an MC scenario in which a macro BS of a cellular network is out of service and multiple UAV-BSs are deployed using integrated access and backhaul (IAB) technology to provide coverage for users in the disaster area. With the data collected from the system-level simulation, a deep reinforcement learning algorithm is developed to jointly optimize the three-dimensional placement of these multiple UAV-BSs, which adapt their 3-D locations to the on-ground user movement. The evaluation results show that the proposed algorithm can support the autonomous navigation of the UAV-BSs to meet the MC service requirements in terms of user throughput and drop rate.
研究动机与目标
- 为解决在传统基站失效的灾难场景中,维持关键任务用户可靠、高吞吐量连接的挑战。
- 在用户移动驱动的动态环境中,联合优化三维UAV-BS定位与回传链路管理。
- 开发一种可扩展的去中心化强化学习框架,实现无需集中协调的UAV-BS自主导航。
- 通过使UAV-BS能够自适应应对用户分布与网络条件的变化,提升应急场景下的系统韧性与服务质量。
提出的方法
- 使用系统级仿真器构建基于5G NR的关键任务场景,包含受损的宏基站与多个采用集成接入与回传(IAB)技术的UAV-BS。
- 提出一种新型去中心化深度强化学习(DecRL)算法——DecRL-AE&VAS,结合自适应探索与基于价值的动作选择策略。
- 通过共享经验回放缓冲区与模型共享机制,提升样本效率,并实现在UAV-BS智能体之间知识的迁移。
- 状态空间包含用户位置、吞吐量与丢弃率;动作空间控制UAV-BS的三维位置(x, y, z);奖励函数为六个性能指标的加权和。
- 采用双轻量双DQN架构估算Q值并指导动作选择,提升学习稳定性与收敛性。
- 通过去中心化决策实现多智能体协同,支持在实时动态环境中大规模部署多个UAV-BS。
实验结果
研究问题
- RQ1在动态关键任务环境中,如何实现多个UAV-BS的自主高效重定位,以维持高用户吞吐量与低丢弃率?
- RQ2在用户移动条件下,去中心化强化学习方法在优化UAV-BS三维部署方面,相较于集中式或基线RL模型,性能提升程度如何?
- RQ3所提出的自适应探索与基于价值的动作选择策略,在变化环境中对提升学习效率与收敛速度的有效性如何?
- RQ4通过复用历史经验,DecRL-AE&VAS算法能否在极少训练迭代下实现近似最优性能?
- RQ5与无UAV或配置欠佳的情况相比,UAV-BS部署对系统性能的影响如何?
主要发现
- DecRL-AE&VAS算法在每次验证阶段的奖励值均达到全局最优解的5%–6%以内,表现出近似最优性能。
- 在宏基站失效后,部署三个UAV-BS使系统性能相比无UAV-BS场景提升约80%,尤其在下行与上行吞吐量及丢弃率方面表现显著。
- 所提方法实现用户吞吐量达10 Mbps,且丢弃率仅比最优解高出2%–3%,表明性能高度一致。
- 通过复用历史经验与基于价值的动作选择,该算法减少了所需训练迭代次数,显著提升学习效率。
- 去中心化架构支持无需集中协调的大规模UAV-BS部署,具备工业级应用潜力。
- 与单个宏基站相比,由三个UAV-BS服务的关键任务用户上行链路5%速率得到提升,归因于更优的空间分布与更近的覆盖距离。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。