Skip to main content
QUICK REVIEW

[论文解读] Intelligent Reflecting Surface Assisted Anti-Jamming Communications: A Fast Reinforcement Learning Approach

Helin Yang, Zehui Xiong|arXiv (Cornell University)|Apr 27, 2020
Advanced Wireless Communication Technologies参考文献 35被引用 10
一句话总结

本文提出一种基于智能反射面(IRS)的快速强化学习方法,以提升无线通信中的抗干扰性能。通过联合优化基站功率分配与IRS反射波束成形,采用模糊胜者为王快速策略爬坡(WoLF-CPHC)算法,该方法在较大IRS元素数量(M=100)下,实现了更高的系统速率与100%的SINR保护水平,优于基线方法,且无需事先掌握干扰模型知识。

ABSTRACT

Malicious jamming launched by smart jammers can attack legitimate transmissions, which has been regarded as one of the critical security challenges in wireless communications. With this focus, this paper considers the use of an intelligent reflecting surface (IRS) to enhance anti-jamming communication performance and mitigate jamming interference by adjusting the surface reflecting elements at the IRS. Aiming to enhance the communication performance against a smart jammer, an optimization problem for jointly optimizing power allocation at the base station (BS), and reflecting beamforming at the IRS is formulated while considering quality of service (QoS) requirements of legitimate users. As the jamming model and jamming behavior are dynamic and unknown, a fuzzy win or learn fast-policy hill-climbing (WoLFPHC) learning approach is proposed to jointly optimize the anti-jamming power allocation and reflecting beamforming strategy, where WoLFPHC is capable of quickly achieving the optimal policy without the knowledge of the jamming model, and fuzzy state aggregation can represent the uncertain environment states as aggregate states. Simulation results demonstrate that the proposed anti-jamming learning-based approach can efficiently improve both the IRS-assisted system rate and transmission protection level compared with existing solutions.

研究动机与目标

  • 解决在未知且动态变化的干扰行为下无线通信中的抗干扰挑战。
  • 通过利用智能反射面(IRS)增强信号并抑制干扰,提升系统速率与传输可靠性。
  • 在服务质量(QoS)约束下,联合优化基站功率分配与IRS反射波束成形。
  • 开发一种快速、无模型的强化学习算法,能够在缺乏干扰特征先验知识的情况下,适应不确定的干扰环境。
  • 通过一种新颖的模糊状态聚合与胜者为王快速策略爬坡(WoLF-CPHC)强化学习框架,提升对智能干扰者的鲁棒性。

提出的方法

  • 建立一个非凸优化问题,以在QoS约束下联合优化基站(BS)的发射功率与IRS的反射波束成形。
  • 提出一种模糊胜者为王快速策略爬坡(WoLF-CPHC)强化学习算法,以在不掌握干扰模型先验知识的情况下,学习最优抗干扰策略。
  • 采用模糊状态聚合(FSA)将不确定的环境状态表示为聚合状态,提升在动态环境中的学习效率。
  • 算法采用SINR感知的奖励函数,优先保障抗干扰能力,提升系统鲁棒性。
  • 通过快速策略收敛,实现实时适应未知干扰行为。
  • IRS通过调整相位偏移,增强期望信号功率并抑制干扰信号,同时功率分配确保满足QoS要求。

实验结果

研究问题

  • RQ1在存在未知智能干扰者的情况下,IRS辅助的波束成形与功率分配如何联合提升抗干扰性能?
  • RQ2无模型强化学习方法是否能在动态、不确定的干扰环境中实现快速收敛与最优策略?
  • RQ3IRS元素数量(M)对抗干扰通信中的系统速率与SINR保护水平有何影响?
  • RQ4所提出的模糊WoLF-CPHC方法与传统Q-learning及基线方法相比,在性能与鲁棒性方面表现如何?
  • RQ5功率分配与反射波束成形的联合优化在多大程度上增强了系统对智能干扰的抗性?

主要发现

  • 当M=100个反射单元时,所提方法实现了12.21 bits/s/Hz的系统速率,显著优于基线方法。
  • 当M≥60时,所提学习方法实现了100%的SINR保护水平,而其他方法表现不足,尤其在高SINR目标下更为明显。
  • 随着IRS元素数量(M)的增加,所提方法与基线方法之间的性能差距进一步扩大,表明其具备良好的可扩展性与适应性。
  • 对于高SINR目标要求(SINR^min > 阈值),系统速率与保护水平因干扰信号而急剧下降,但IRS辅助方法仍保持优异性能。
  • 在系统速率与SINR保护方面,基于模糊WoLF-CPHC的方法均优于快速Q-learning与基线方法,尤其在严格的QoS约束下表现更优。
  • 联合优化功率分配与反射波束成形对最大化抗干扰性能至关重要,如所提方法相较其他方案的优越表现所示。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。