Skip to main content
QUICK REVIEW

[论文解读] A General Deep Reinforcement Learning Framework for Grant-Free NOMA Optimization in mURLLC

Yan Liu, Yansha Deng|arXiv (Cornell University)|Jan 2, 2021
Advanced Wireless Communication Technologies参考文献 29被引用 6
一句话总结

本文提出一种通用的深度强化学习框架,结合双重DQN(Double DQN)与合作多智能体DQN(CMA-DQN),用于在大规模超可靠低延迟通信(mURLLC)的免许可非正交多址(grant-free NOMA)系统中优化资源配置。该框架动态调整重复次数与竞争-传输单元(CTU)数量,在相同延迟约束下,相较于固定配置,K-重复GF-NOMA的可成功服务用户数提升最高达10倍,Proactive GF-NOMA提升2倍。

ABSTRACT

Grant-free non-orthogonal multiple access (GF-NOMA) is a potential technique to support massive Ultra-Reliable and Low-Latency Communication (mURLLC) service. However, the dynamic resource configuration in GF-NOMA systems is challenging due to random traffics and collisions, that are unknown at the base station (BS). Meanwhile, joint consideration of the latency and reliability requirements makes the resource configuration of GF-NOMA for mURLLC more complex. To address this problem, we develop a general learning framework for signature-based GF-NOMA in mURLLC service taking into account the multiple access signature collision, the UE detection, as well as the data decoding procedures for the K-repetition GF and the Proactive GF schemes. The goal of our learning framework is to maximize the long-term average number of successfully served users (UEs) under the latency constraint. We first perform a real-time repetition value configuration based on a double deep Q-Network (DDQN) and then propose a Cooperative Multi-Agent learning technique based on the DQN (CMA-DQN) to optimize the configuration of both the repetition values and the contention-transmission unit (CTU) numbers. Our results show that the number of successfully served UEs under the same latency constraint in our proposed learning framework is up to ten times for the K-repetition scheme, and two times for the Proactive scheme, more than that with fixed repetition values and CTU numbers. In addition, the superior performance of CMA-DQN over the conventional load estimation-based approach (LE-URC) demonstrates its capability in dynamically configuring in long term. Importantly, our general learning framework can be used to optimize the resource configuration problems in all the signature-based GF-NOMA schemes.

研究动机与目标

  • 为解决免许可NOMA(GF-NOMA)系统在mURLLC场景下面临的动态资源配置挑战,其中基站无法预知随机流量与碰撞情况。
  • 通过管理导频碰撞、用户设备(UE)检测与数据解码流程,联合优化GF-NOMA中的时延与可靠性需求。
  • 开发一种可泛化的学习框架,适用于所有基于导频的GF-NOMA方案,包括K-重复与Proactive GF。
  • 在严格时延约束下,最大化长期平均成功服务UE数量。
  • 在动态、长期资源配置场景中,优于传统的基于负载估计的方案(LE-URC)

提出的方法

  • 采用双重深度Q网络(DDQN)根据当前系统状态与反馈,实现实时重复次数配置。
  • 提出合作多智能体DQN(CMA-DQN),联合优化多个UE的重复次数与竞争-传输单元(CTU)数量。
  • 将GF-NOMA系统建模为马尔可夫决策过程(MDP),其中状态表示信道条件、负载情况与用户活动状态。
  • DQN智能体通过基于成功用户解码与时延合规性的延迟奖励学习策略。
  • CMA-DQN架构使代表不同UE或资源块的多个智能体实现协调,以最小化碰撞并提升频谱效率。
  • 该框架设计具有通用性与可扩展性,可直接应用于任何基于导频的GF-NOMA方案,包括K-重复与Proactive GF。

实验结果

研究问题

  • RQ1如何利用深度强化学习在mURLLC的GF-NOMA中动态配置重复次数与CTU数量?
  • RQ2CMA-DQN方法在成功服务用户数方面,相较于固定或启发式配置策略,能提升多少?
  • RQ3所提出的框架能否在K-重复与Proactive GF等不同GF-NOMA方案间实现良好泛化?
  • RQ4该学习框架在流量未知且动态变化的条件下,如何维持高可靠性和低时延?
  • RQ5与传统的基于负载估计的资源控制(LE-URC)相比,CMA-DQN方法在长期性能上具有多大提升?

主要发现

  • 在K-重复GF-NOMA方案中,所提出的CMA-DQN框架在相同时延约束下,成功服务用户数最高可达固定重复值配置的10倍。
  • 在Proactive GF-NOMA方案中,该框架相较固定配置方法,成功服务用户数提升2倍。
  • CMA-DQN方法在长期动态资源配置中显著优于传统的基于负载估计的资源控制(LE-URC)方法,展现出更强的适应能力。
  • 该框架在统一的学习驱动优化框架中,有效处理了导频碰撞、UE检测与数据解码流程。
  • 该框架的通用化设计可直接应用于任何基于导频的GF-NOMA方案,确保了广泛的适用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。