Skip to main content
QUICK REVIEW

[论文解读] Two-dimensional Anti-jamming Mobile Communication Based on Reinforcement Learning

Liang Xiao, Guoan Han|arXiv (Cornell University)|Dec 19, 2017
Security in Wireless Sensor Networks参考文献 24被引用 8
一句话总结

本文提出了一种结合跳频和用户移动性的二维抗干扰移动通信系统,采用带有热启动和宏观动作的快速深度Q网络(DQN),使移动设备能够在不了解干扰或信道模型的情况下自主抵抗协同干扰。与贪婪基线相比,该方法实现90%更快的学习速度和84.7%更高的信干噪比(SINR),在动态环境中实现92.1%更高的效用。

ABSTRACT

By using smart radio devices, a jammer can dynamically change its jamming policy based on opposing security mechanisms; it can even induce the mobile device to enter a specific communication mode and then launch the jamming policy accordingly. On the other hand, mobile devices can exploit spread spectrum and user mobility to address both jamming and interference. In this paper, a two-dimensional anti-jamming mobile communication scheme is proposed in which a mobile device leaves a heavily jammed/interfered-with frequency or area. It is shown that, by applying reinforcement learning techniques, a mobile device can achieve an optimal communication policy without the need to know the jamming and interference model and the radio channel model in a dynamic game framework. More specifically, a hotbooting deep Q-network based two-dimensional mobile communication scheme is proposed that exploits experiences in similar scenarios to reduce the exploration time at the beginning of the game, and applies deep convolutional neural network and macro-action techniques to accelerate the learning speed in dynamic situations. Several real-world scenarios are simulated to evaluate the proposed method. These simulation results show that our proposed scheme can improve both the signal-to-interference-plus-noise ratio of the signals and the utility of the mobile devices against cooperative jamming compared with benchmark schemes.

研究动机与目标

  • 解决传统扩频技术在存在协同干扰和强干扰环境下的局限性。
  • 设计一种利用跳频和用户移动性的动态抗干扰策略,以增强系统鲁棒性。
  • 使移动设备能够在不了解干扰、干扰或信道模型的情况下学习最优通信策略。
  • 加速典型真实移动通信场景中高维状态空间的学习速度。
  • 通过深度强化学习提升动态干扰条件下的信号质量和效用。

提出的方法

  • 采用深度Q网络(DQN)对二维抗干扰框架中的功率分配与移动决策建模为马尔可夫决策过程(MDP)。
  • 使用深度卷积神经网络(CNN)压缩并表示高维状态观测,降低状态空间复杂度。
  • 引入宏观动作,将多个时间步的决策(功率与移动)合并为单一动作,提升学习效率。
  • 采用热启动技术,利用相似场景的预训练模型初始化CNN权重,减少初始探索时间。
  • 系统在动态博弈框架中进行训练,移动设备通过与不断演化的干扰源和干扰进行试错交互学习。
  • 在三种真实场景中评估该方法:指令分发、传感报告传输和移动干扰抵抗。

实验结果

研究问题

  • RQ1基于强化学习的系统是否能在不了解干扰或信道模型的情况下,有效学习最优抗干扰策略?
  • RQ2结合跳频与用户移动性在多大程度上提升了对协同干扰和干扰的鲁棒性?
  • RQ3热启动与宏观动作的集成在高维2D抗干扰环境中在多大程度上加速了学习?
  • RQ4与标准DQN、Q-learning和贪婪基线相比,所提出的快速DQN方案在性能和学习速度上表现如何?
  • RQ5该系统对动态移动的干扰源具有多强的鲁棒性?

主要发现

  • 与标准DQN相比,快速DQN方案将学习时间减少90%,显著加速了动态抗干扰博弈中的收敛速度。
  • 在64个信道系统中,快速DQN方案相比基于贪婪的方案实现84.7%更高的信干噪比(SINR)。
  • 与基于贪婪的方案相比,快速DQN方案使效用提升92.1%,展现出更优的长期性能。
  • 在64个信道条件下,快速DQN方案相比DQN方案实现12.8%更高的SINR,相比贪婪方案实现72.5%更高的SINR。
  • 在移动干扰源条件下,系统保持强性能,当干扰源每200个时间槽随机移动一次时,SINR仅下降0.6%,效用仅下降1.1%。
  • 当单位传输成本从0.1增加到0.3时,DQN方案的SINR下降4.9%,效用下降63.3%,但快速DQN方案在成本约束下仍保持优越性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。