[论文解读] Reinforcement Learning for Interference Avoidance Game in RF-Powered Backscatter Communications
本文提出了一种基于强化学习的射频供能背向散射通信中的干扰规避策略,将用户与智能干扰器之间的交互建模为斯塔克尔伯格博弈。通过将效用定义为传输比特数,并使用带有热启动(hotbooting)的Q-learning以加快收敛速度,该方法相比标准Q-learning将用户效用提高了最多31%,并在动态、不确定环境中优于随机和固定策略。
RF-powered backscatter communication is a promising new technology that can be deployed for battery-free applications such as internet of things (IoT) and wireless sensor networks (WSN). However, since this kind of communication is based on the ambient RF signals and battery-free devices, they are vulnerable to interference and jamming. In this paper, we model the interaction between the user and a smart interferer in an ambient backscatter communication network as a game. We design the utility functions of both the user and interferer in which the backscattering time is taken into the account. The convexity of both sub-game optimization problems is proved and the closed-form expression for the equilibrium of the Stackelberg game is obtained. Due to lack of information about the system SNR and transmission strategy of the interferer, the optimal strategy is obtained using the Q-learning algorithm in a dynamic iterative manner. We further introduce hotbooting Q-learning as an effective approach to expedite the convergence of the traditional Q-learning. Simulation results show that our approach can obtain considerable performance improvement in comparison to random and fixed backscattering time transmission strategies and improves the convergence speed of Q-Learning by about 31%.
研究动机与目标
- 为解决在不确定环境中用户与智能干扰器动态交互时的干扰问题,提出在射频供能背向散射通信中的干扰挑战。
- 将背向散射时延和传输功率分配建模为以传输比特数为效用定义的斯塔克尔伯格博弈。
- 为用户开发一种基于学习的策略,以在缺乏对系统状态和干扰器行为了解的情况下自适应选择背向散射时延。
- 通过一种新颖的热启动技术加速Q-learning的收敛速度,实现在动态场景中的更快策略学习。
提出的方法
- 将用户与智能干扰器之间的交互建模为斯塔克尔伯格博弈,其中用户为领导者,干扰器为跟随者。
- 基于传输比特数定义效用函数,结合背向散射时延和传输功率成本。
- 证明子博弈优化问题的凸性,并通过解析方法推导出闭式均衡解。
- 在系统状态和干扰器策略未知的动态、部分可观察环境中应用Q-learning。
- 通过使用解析均衡解初始化Q值,实现热启动Q-learning,以加速收敛。
- 为背向散射时延和干扰器功率采用离散动作空间,结合ε-greedy探索和时序差分更新。
实验结果
研究问题
- RQ1在射频供能背向散射网络中,面对一个智能且自适应的干扰器,用户如何最优地选择其背向散射时延?
- RQ2在双方对彼此行为均不完全了解的动态斯塔克尔伯格博弈中,其均衡策略是什么?
- RQ3Q-learning能否在部分可观察、易受干扰的背向散射环境中有效学习最优传输策略?
- RQ4与标准Q-learning相比,热启动Q-learning在该动态博弈设置中在多大程度上提升了收敛速度?
主要发现
- 所提出的基于Q-learning的方法相比随机和固定背向散射时延策略,显著提升了用户效用。
- 在仿真设置中,热启动Q-learning相比标准Q-learning将收敛速度加快了约31%。
- 随着用户与高空平台站(HAP)之间距离的增加,用户平均效用下降,反映出路径损耗对能量和数据采集的影响。
- 当背向散射时延的成本增加时,用户效用随之降低,表明传输与能量采集之间存在权衡。
- 标准Q-learning与热启动Q-learning均达到相似的最终效用水平,证实其收敛至近似最优策略。
- 斯塔克尔伯格博弈均衡的解析解有效且可作为学习过程的优良初始化。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。