Skip to main content
QUICK REVIEW

[论文解读] AI-based Radio Resource Management and Trajectory Design for PD-NOMA Communication in IRS-UAV Assisted Networks

Hussein M. Hariz, Saeed Sheikhzadeh|arXiv (Cornell University)|Nov 6, 2021
Age of Information Optimization被引用 10
一句话总结

本文提出了一种基于人工智能的无人机-智能反射面(UAV-IRS)辅助物理层非正交多址(PD-NOMA)网络中的无线资源管理与轨迹优化框架,旨在最小化物联网(IoT)系统中的信息年龄(AAoI)。通过结合双重深度Q网络(DDQN)与近端策略优化(PPO),该方法联合优化了功率、子载波、轨迹及相位移位,相较于匹配算法和随机轨迹基线,AAoI性能分别提升了10%和15%。

ABSTRACT

In this paper, we consider that the unmanned aerial vehicles (UAVs) with attached intelligent reflecting surfaces (IRSs) play the role of flying reflectors that reflect the signal of users to the destination, and utilize the power-domain non-orthogonal multiple access (PD-NOMA) scheme in the uplink. We investigate the benefits of the UAV-IRS on the internet of things (IoT) networks that improve the freshness of collected data of the IoT devices via optimizing power, sub-carrier, and trajectory variables, as well as, the phase shift matrix elements. We consider minimizing the average age-of-information (AAoI) of users subject to the maximum transmit power limitations, PD-NOMA-related restriction, and the constraints related to UAV's movement. The optimization problem consists of discrete and continuous variables. Hence, we divide the resource allocation problem into two sub-problems and use two different reinforcement learning (RL) based algorithms to solve them, namely the double deep Qnetwork (DDQN) and a proximal policy optimization (PPO). Our numerical results illustrate the performance gains that can be achieved for IRS enabled UAV communication systems. Moreover, we compare our deep RL (DRL) based algorithm with matching algorithm and random trajectory, showing the combination of DDQN and PPO algorithm proposed in this paper performs 10% and 15% better than matching algorithm and random-trajectory algorithm, respectively.

研究动机与目标

  • 解决具有高数据新鲜度要求的物联网网络中及时状态更新传输的挑战。
  • 在实际约束条件下,最小化UAV-IRS辅助PD-NOMA上行链路通信中的平均信息年龄(AAoI)。
  • 同时优化功率分配、子载波分配、无人机(UAV)轨迹以及智能反射面(IRS)相位移位。
  • 利用深度强化学习克服动态无线环境中混合整数非凸优化的复杂性。
  • 展示所提出的深度强化学习(DRL)框架相较于传统基线方法(如匹配算法与随机轨迹算法)的性能增益。

提出的方法

  • 建立一个混合整数非凸优化问题,以在功率、PD-NOMA及UAV移动性约束下最小化AAoI。
  • 将问题分解为两个子问题:子载波与功率分配(通过DDQN求解),以及UAV轨迹设计(通过PPO求解)。
  • 使用双重深度Q网络(DDQN)智能体,基于反映AAoI的状态相关奖励学习最优的功率与子载波分配。
  • 采用近端策略优化(PPO)智能体,通过连续动作空间控制学习最优UAV轨迹,以降低AAoI。
  • 将IRS相位移位作为动态波束成形组件,增强远距离用户的信号强度与可靠性。
  • 在模拟环境中端到端训练两个智能体,其状态表示包括用户位置、信道条件及AAoI历史。

实验结果

研究问题

  • RQ1如何优化UAV-IRS辅助的PD-NOMA系统,以最小化物联网网络中的平均信息年龄(AAoI)?
  • RQ2与独立无人机系统相比,将智能反射面(IRS)与无人机集成在多大程度上能提升AAoI性能?
  • RQ3DDQN与PPO的结合在联合优化功率、子载波、轨迹及相位移位以实现AAoI最小化方面有多高效?
  • RQ4与基线方案(如匹配算法与随机轨迹)相比,所提出的DRL框架可实现多大的性能增益?
  • RQ5发射功率限制如何影响AAoI的降低?在何种场景下部署IRS最具优势?

主要发现

  • 所提出的DDQN-PPO框架相较于匹配算法基线,AAoI性能提升了10%。
  • 与随机轨迹基线相比,该框架使AAoI降低了15%,证明了智能轨迹设计的有效性。
  • DRL智能体在4,000个训练周期内实现收敛,学习过程稳定,负奖励趋势表明AAoI随时间持续改善。
  • 由PPO学习到的UAV轨迹自适应地朝向AAoI较高且信道条件较弱的用户移动,从而提升数据速率并减少信息年龄。
  • 在低发射功率约束下,IRS部署显著提升了系统性能,使其成为功率受限物联网应用的关键使能技术。
  • 随着用户数量增加,AAoI因频谱与时间资源受限而上升,但所提方法在所有用户数量下均保持低于基线的AAoI水平。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。