[论文解读] Competitive MA-DRL for Transmit Power Pool Design in Semi-Grant-Free NOMA Systems.
本文提出了一种竞争性多智能体深度强化学习(MA-DRL)框架,用于在大规模物联网网络中为半无许可免授权非正交多址接入(SGF-NOMA)设计动态发射功率池(PP)。通过将资源选择建模为随机马尔可夫博弈,并采用双值Dueling DDQN以提升学习效率,该方法相较于纯无许可免授权协议实现了22.2%的系统吞吐量增益,相较于固定功率控制方案提升了17.5%,同时通过无效动作剪枝显著缩短了训练时间。
In this paper, we exploit the capability of multi-agent deep reinforcement learning (MA-DRL) technique to generate a transmit power pool (PP) for Internet of things (IoT) networks with semi-grant-free non-orthogonal multiple access (SGF-NOMA). The PP is mapped with each resource block (RB) to achieve distributed transmit power control (DPC). We first formulate the resource (sub-channel and transmit power) selection problem as stochastic Markov game, and then solve it using two competitive MA-DRL algorithms, namely double deep Q network (DDQN) and Dueling DDQN. Each GF user as an agent tries to find out the optimal transmit power level and RB to form the desired PP. With the aid of dueling processes, the learning process can be enhanced by evaluating the valuable state without considering the effect of each action at each state. Therefore, DDQN is designed for communication scenarios with a small-size action-state space, while Dueling DDQN is for a large-size case. Our results show that the proposed MA-Dueling DDQN based SGF-NOMA with DPC outperforms the SGF-NOMA system with the fixed-power-control mechanism and networks with pure GF protocols with 17.5% and 22.2% gain in terms of the system throughput, respectively. Moreover, to decrease the training time, we eliminate invalid actions (high transmit power levels) to reduce the action space. We show that our proposed algorithm is computationally scalable to massive IoT networks. Finally, to control the interference and guarantee the quality-of-service requirements of grant-based users, we find the optimal number of GF users for each sub-channel.
研究动机与目标
- 通过实现分布式发射功率控制,解决大规模物联网网络中半无许可免授权NOMA的高效资源分配挑战。
- 克服固定功率控制和纯无许可免授权协议在维持服务质量与频谱效率方面的局限性。
- 设计一种可扩展、低时延的功率池机制,支持大规模连接与干扰管理。
- 优化每个子信道上的无许可用户数量,以在吞吐量与有许可用户的服务质量之间实现平衡。
- 通过双值网络架构与无效动作消除,提升多智能体环境中的学习效率。
提出的方法
- 将联合子信道与发射功率选择问题建模为随机马尔可夫博弈,以刻画无许可用户之间的交互行为。
- 实现两种竞争性MA-DRL算法——DDQN与Dueling DDQN,其中每个无许可用户作为独立智能体,自主优化其功率与子信道选择。
- 引入双值网络结构,解耦状态值与优势估计,从而提升学习稳定性与收敛速度。
- 通过消除高发射功率水平等无效或不切实际的动作,缩小动作空间,从而加速训练并提升可扩展性。
- 将所得最优功率水平映射为每个资源块的发射功率池(PP),以实现网络范围内的分布式功率控制。
- 确定每个子信道上最优的无许可用户数量,以最小化干扰并确保有许可用户的QoS。
实验结果
研究问题
- RQ1多智能体深度强化学习如何被有效应用于半无许可免授权NOMA物联网网络中动态发射功率池的设计?
- RQ2在大规模动作-状态空间中,使用Dueling DDQN相较于标准DDQN对学习效率与系统吞吐量有何影响?
- RQ3在大规模物联网部署中,剪枝无效动作在多大程度上可提升训练时间与可扩展性?
- RQ4所提出的基于MA-DRL的功率池设计相较于固定功率控制与纯无许可免授权协议,在系统吞吐量方面表现如何?
- RQ5每个子信道上最优的无许可用户数量是多少,才能在维持有许可用户QoS的同时最大化系统吞吐量?
主要发现
- 所提出的基于MA-Dueling DDQN的SGF-NOMA与分布式功率控制方案,相较于纯无许可免授权协议,实现了22.2%更高的系统吞吐量。
- 在系统吞吐量方面,该系统相较于固定功率控制机制提升了17.5%,证明了自适应功率分配的优势。
- 剪枝无效的高功率动作显著缩小了有效动作空间,从而加快了训练收敛速度并提升了计算可扩展性。
- 双值网络结构通过解耦状态值与优势估计,在大规模动作-状态环境中显著提升了学习稳定性与性能表现。
- 确定了每个子信道上最优的无许可用户数量,以在频谱效率与干扰之间实现平衡,同时保障有许可用户的QoS。
- 整体框架具有良好的计算可扩展性,适用于高用户密度的大规模物联网网络部署。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。