[论文解读] Distributed 3D-Beam Reforming for Hovering-Tolerant UAVs Communication over Coexistence: A Deep-Q Learning for Intelligent Space-Air-Ground Integrated Networks
该论文提出了一种无模型、基于深度Q-learning的分布式3D波束成形框架,适用于在空间-空中-地面一体化网络中具有悬停容忍度的无人机(UAV),可在动态运动和共存干扰条件下实现实时波束跟踪与重构。该方法在极低均方误差下实现快速收敛(≈50次迭代),即使在无人机随机旋转和悬停引起的波束失准情况下,仍能保持高 spectral efficiency 和 QoS。
In this paper, we present a novel distributed UAVs beam reforming approach to dynamically form and reform a space-selective beam path in addressing the coexistence with satellite and terrestrial communications. Despite the unique advantage to support wider coverage in UAV-enabled cellular communications, the challenges reside in the array responses' sensitivity to random rotational motion and the hovering nature of the UAVs. A model-free reinforcement learning (RL) based unified UAV beam selection and tracking approach is presented to effectively realize the dynamic distributed and collaborative beamforming. The combined impact of the UAVs' hovering and rotational motions is considered while addressing the impairment due to the interference from the orbiting satellites and neighboring networks. The main objectives of this work are two-fold: first, to acquire the channel awareness to uncover its impairments; second, to overcome the beam distortion to meet the quality of service (QoS) requirements. To overcome the impact of the interference and to maximize the beamforming gain, we define and apply a new optimal UAV selection algorithm based on the brute force criteria. Results demonstrate that the detrimental effects of the channel fading and the interference from the orbiting satellites and neighboring networks can be overcome using the proposed approach. Subsequently, an RL algorithm based on Deep Q-Network (DQN) is developed for real-time beam tracking. By augmenting the system with the impairments due to hovering and rotational motion, we show that the proposed DQN algorithm can reform the beam in real-time with negligible error. It is demonstrated that the proposed DQN algorithm attains an exceptional performance improvement. We show that it requires a few iterations only for fine-tuning its parameters without observing any plateaus irrespective of the hovering tolerance.
研究动机与目标
- 解决无人机在三维空间-空中-地面一体化网络中因随机旋转和悬停运动引起的波束失准问题。
- 在保持高频谱效率和服务质量(QoS)的同时,克服轨道卫星和相邻地面网络的干扰。
- 开发一种无模型的强化学习框架,实现实时波束选择与跟踪,无需依赖信道状态信息或复杂预测模型。
- 在动态环境与移动性约束下,实现无人机之间的协作式、分布式波束成形,动态形成最优波束路径。
- 确保在不同悬停容忍度水平下均具备鲁棒性能,表现出对运动引起的畸变的独立性。
提出的方法
- 提出一种暴力搜索最优UAV选择算法,从N=64架UAV中基于方向和信道条件选出最佳K架UAV(例如K=4),以最大化SINR并最小化时延。
- 引入一种深度Q网络(DQN)智能体,采用深度神经网络(DNN)作为函数逼近器,实时估计波束成形动作的Q值。
- 采用经验回放技术以稳定DQN训练过程中的学习并提高样本效率。
- 将无人机悬停和旋转运动(偏航、俯仰、滚转±10°)的影响建模为波束张力和失准的来源,并将这些损伤纳入环境动力学中。
- 采用64架UAV的三维矩形UAV构型,并模拟来自随机分布的相邻网络和轨道卫星的干扰,以反映真实世界的共存挑战。
- 集成信道感知机制以检测干扰并识别最优链路以实现机会访问,从而支持动态波束重构。

实验结果
研究问题
- RQ1无人机随机悬停和旋转运动如何影响三维空间-空中-地面一体化网络中的波束成形增益和波束对准?
- RQ2在动态运动和干扰条件下,何种UAV选择策略可最大化接收SINR并最小化时延?
- RQ3无模型的深度强化学习方法是否能有效实现实时波束跟踪与重构,即使在无人机移动引起的波束失准情况下?
- RQ4所提出的基于DQN的波束成形算法在不同悬停容忍度水平下,其收敛速度和稳定性表现如何?
- RQ5所提方法在保持QoS和频谱效率的同时,能在多大程度上缓解共存卫星和地面网络的干扰?
主要发现
- 所提出的暴力搜索UAV选择算法能有效从64架UAV中识别出最优的4架UAV,在干扰条件下最大化SINR并最小化时延。
- 基于DQN的波束跟踪算法收敛迅速,仅需约50次迭代即可微调参数,且未观察到性能平台期。
- 波束重构技术实现了可忽略的均方误差(MSE),表明其在纠正由无人机运动引起的波束失准方面具有高精度。
- 学习算法在不同悬停容忍度值下均表现出高效性且独立运行,即使在相邻UAV间距的30%(δ = 1 m)时仍保持鲁棒性。
- 增加UAV间距可提升波束方向性和干扰抑制能力,但同时也会提高旁瓣电平,凸显阵列设计中的权衡。
- 该系统能有效缓解来自轨道卫星和相邻网络的干扰,在动态条件下保持高谱效率和QoS。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。