Skip to main content
QUICK REVIEW

[论文解读] Multi-Agent Reinforcement Learning based Joint Cooperative Spectrum Sensing and Channel Access for Cognitive UAV Networks.

Weiheng Jiang, Wanxin Yu|arXiv (Cornell University)|Mar 15, 2021
Cognitive Radio Networks and Spectrum Sensing参考文献 44被引用 5
一句话总结

本文提出了一种多智能体强化学习(MARL)框架,用于认知无人机(UAV)网络中的联合协作式频谱感知与信道接入,采用加权奖励函数以平衡感知成本与传输效用。IL-Q、IL-DQN 和 VC-EXH 算法在动态、非平稳环境中实现了快速收敛,并显著提升了频谱利用率。

ABSTRACT

Designing clustered unmanned aerial vehicle (UAV) communication networks based on cognitive radio (CR) and reinforcement learning can significantly improve the intelligence level of clustered UAV communication networks and the robustness of the system in a time-varying environment. Among them, designing smarter systems for spectrum sensing and access is a key research issue in CR. Therefore, we focus on the dynamic cooperative spectrum sensing and channel access in clustered cognitive UAV (CUAV) communication networks. Due to the lack of prior statistical information on the primary user (PU) channel occupancy state, we propose to use multi-agent reinforcement learning (MARL) to model CUAV spectrum competition and cooperative decision-making problem in this dynamic scenario, and a return function based on the weighted compound of sensing-transmission cost and utility is introduced to characterize the real-time rewards of multi-agent game. On this basis, a time slot multi-round revisit exhaustive search algorithm based on virtual controller (VC-EXH), a Q-learning algorithm based on independent learner (IL-Q) and a deep Q-learning algorithm based on independent learner (IL-DQN) are respectively proposed. Further, the information exchange overhead, execution complexity and convergence of the three algorithms are briefly analyzed. Through the numerical simulation analysis, all three algorithms can converge quickly, significantly improve system performance and increase the utilization of idle spectrum resources.

研究动机与目标

  • 解决在缺乏主用户活动先验知识的簇状认知无人机网络中动态频谱接入的挑战。
  • 通过无人机之间的自主协作决策,提升时变无线环境下的系统智能与鲁棒性。
  • 优化频谱感知与接入决策中感知成本与传输效用之间的权衡。
  • 设计可扩展且收敛的算法,以减少信息交换开销与执行复杂度。
  • 通过多无人机智能体之间的联合感知与接入协调,提升空闲频谱资源的利用率。

提出的方法

  • 将频谱感知与接入问题建模为一个多智能体强化学习博弈,其中无人机作为独立的学习者。
  • 引入一种结合感知成本与传输效用的加权复合奖励函数,以反映实时性能增益。
  • 提出一种基于虚拟控制器的时间槽多轮回溯遍历搜索算法(VC-EXH),用于协调智能体决策。
  • 设计一种独立 Q-learning(IL-Q)算法,实现基于表格 Q 值学习的去中心化决策。
  • 设计一种独立深度 Q-network(IL-DQN),以处理高维状态空间并提升学习稳定性。
  • 分析算法复杂度、收敛速度与通信开销,以评估其可扩展性与实用性。

实验结果

研究问题

  • RQ1多智能体强化学习如何在动态、非平稳的认知无人机网络中有效协调频谱感知与接入?
  • RQ2在协作式频谱感知中,何种奖励函数设计能最佳平衡感知成本与传输效用?
  • RQ3IL-Q、IL-DQN 与 VC-EXH 在收敛速度、系统性能与通信开销方面有何比较差异?
  • RQ4基于 MARL 的方法是否能在缺乏主用户行为统计知识的前提下显著提升频谱利用率?
  • RQ5在去中心化无人机频谱接入中,算法复杂度、信息交换与收敛性之间存在何种权衡?

主要发现

  • 所提出的三种算法——VC-EXH、IL-Q 与 IL-DQN——在动态频谱接入场景中均实现了快速收敛。
  • 基于 MARL 的方法通过提升空闲频谱资源的利用率,显著改善了系统性能。
  • 加权复合奖励函数有效平衡了感知成本与传输效用,从而实现更优的实时决策。
  • 与基于表格的 IL-Q 相比,IL-DQN 在高维状态空间中表现出更优的学习稳定性。
  • VC-EXH 通过使用虚拟控制器协调智能体行为,避免了完全通信,从而降低了信息交换开销。
  • 这些算法共同降低了系统时延,并增强了在时变无线环境下的鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。