[论文解读] Machine Learning Empowered Trajectory and Passive Beamforming Design in UAV-RIS Wireless Networks
该论文提出了一种深度强化学习框架,通过使用衰减深度Q网络(D-DQN)联合优化无人机轨迹、智能反射面(RIS)相位移位、功率分配和动态解码顺序,以最小化无人机能耗。结果表明,与RIS-OMA相比,RIS-NOMA可将能耗降低11.7%,且D-DQN算法成功收敛,而传统Q-learning则未能收敛。
A novel framework is proposed for integrating reconfigurable intelligent surfaces (RIS) in unmanned aerial vehicle (UAV) enabled wireless networks, where an RIS is deployed for enhancing the service quality of the UAV. Non-orthogonal multiple access (NOMA) technique is invoked to further improve the spectrum efficiency of the network, while mobile users (MUs) are considered as roaming continuously. The energy consumption minimizing problem is formulated by jointly designing the movement of the UAV, phase shifts of the RIS, power allocation policy from the UAV to MUs, as well as determining the dynamic decoding order. A decaying deep Q-network (D-DQN) based algorithm is proposed for tackling this pertinent problem. In the proposed D-DQN based algorithm, the central controller is selected as an agent for periodically observing the state of UAV-enabled wireless network and for carrying out actions to adapt to the dynamic environment. In contrast to the conventional DQN algorithm, the decaying learning rate is leveraged in the proposed D-DQN based algorithm for attaining a tradeoff between accelerating training speed and converging to the local optimal. Numerical results demonstrate that: 1) In contrast to the conventional Q-learning algorithm, which cannot converge when being adopted for solving the formulated problem, the proposed D-DQN based algorithm is capable of converging with minor constraints; 2) The energy dissipation of the UAV can be significantly reduced by integrating RISs in UAV-enabled wireless networks; 3) By designing the dynamic decoding order and power allocation policy, the RIS-NOMA case consumes 11.7% less energy than the RIS-OMA case.
研究动机与目标
- 解决无人机增强型无线网络中的高能耗与动态环境挑战。
- 联合优化无人机轨迹、RIS相位移位、功率分配与解码顺序以实现能耗最小化。
- 实现对移动用户移动性与信道变化的动态、实时自适应。
- 利用RIS与NOMA提升频谱效率并扩展覆盖范围。
- 开发一种可在复杂、高维状态空间中可靠收敛的基于学习的解决方案。
提出的方法
- 采用衰减深度Q网络(D-DQN),无人机的中心控制器作为智能体,观测网络状态并采取行动。
- D-DQN引入衰减学习率,以平衡训练速度与局部最优解的收敛性。
- 状态空间包括无人机位置、用户位置、信道状态信息及用户数据需求。
- 动作包括无人机移动指令、RIS相位移位调整、功率分配水平以及NOMA中的动态解码顺序选择。
- 奖励函数设计为惩罚能耗并奖励成功满足用户速率约束的数据传输。
- 该算法在统一的强化学习框架下联合优化轨迹、波束成形、功率分配与解码顺序。
实验结果
研究问题
- RQ1基于D-DQN的方法能否在无人机-RIS轨迹与波束成形设计的复杂、高维优化问题中实现有效收敛?
- RQ2与传统的RIS-OMA或非RIS系统相比,集成RIS与NOMA对无人机能耗有何影响?
- RQ3在无人机-NOMA-RIS网络中,动态解码顺序与自适应功率分配对能效有何影响?
- RQ4RIS反射单元数量与无人机高度如何影响能耗与系统性能?
- RQ5D-DQN算法在收敛性与性能方面相较于传统Q-learning的优越程度如何?
主要发现
- 所提出的D-DQN算法成功收敛至稳定策略,而传统Q-learning在相同问题设定下无法收敛。
- 在无人机网络中集成RIS显著降低了无人机能耗,通过被动反射实现更可靠的视 Line-of-Sight(LoS)链路。
- 由于频谱效率更高且减少了无人机移动需求,RIS-NOMA相比RIS-OMA将能耗降低了11.7%。
- 增加RIS反射单元数量可降低无人机能耗,当服务三个用户簇时能耗相比两个用户簇增加10.3%。
- 与固定策略相比,动态解码顺序与自适应功率分配可降低能耗,因其在用户移动条件下维持了更高的频谱效率。
- 无人机高度对能耗具有非单调影响:200米高度时能耗更高,尽管视 Line-of-Sight(LoS)概率更高,但路径损耗增加。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。