[论文解读] Machine Learning for User Partitioning and Phase Shifters Design in RIS-Aided NOMA Networks
本文提出了一种基于机器学习的框架,用于在RIS辅助的NOMA网络中联合进行用户分组和相位移位优化,采用改进的MOMA算法进行聚类,以及基于DDPG的强化学习方法实现动态RIS波束成形。与传统的OMA和随机相位移位策略相比,该方法在频谱效率增益方面表现显著,且随着RIS相位移位分辨率提高和反射单元数量增加,性能进一步增强。
A novel reconfigurable intelligent surface (RIS) aided non-orthogonal multiple access (NOMA) downlink transmission framework is proposed. We formulate a long-term stochastic optimization problem that involves a joint optimization of NOMA user partitioning and RIS phase shifting, aiming at maximizing the sum data rate of the mobile users (MUs) in NOMA downlink networks. To solve the challenging joint optimization problem, we invoke a modified object migration automation (MOMA) algorithm to partition the users into equal-size clusters. To optimize the RIS phase-shifting matrix, we propose a deep deterministic policy gradient (DDPG) algorithm to collaboratively control multiple reflecting elements (REs) of the RIS. Different from conventional training-then-testing processing, we consider a long-term self-adjusting learning model where the intelligent agent is capable of learning the optimal action for every given state through exploration and exploitation. Extensive numerical results demonstrate that: 1) The proposed RIS-aided NOMA downlink framework achieves an enhanced sum data rate compared with the conventional orthogonal multiple access (OMA) framework. 2) The proposed DDPG algorithm is capable of learning a dynamic resource allocation policy in a long-term manner. 3) The performance of the proposed RIS-aided NOMA framework can be improved by increasing the granularity of the RIS phase shifts. The numerical results also show that reducing the granularity of the RIS phase shifts and increasing the number of REs are two efficient methods to improve the sum data rate of the MUs.
研究动机与目标
- 解决在RIS辅助的NOMA下行链路网络中,通过联合用户分组与RIS相位移位优化以最大化总数据速率的挑战。
- 克服在具有移动用户的动态无线环境中,随机性长期优化的复杂性。
- 开发一种自适应学习模型,使智能体能够随时间探索并利用最优动作。
- 评估RIS相位移位粒度和反射单元数量对系统性能的影响。
提出的方法
- 采用改进的物体迁移自动化(MOMA)算法,将用户划分为大小相等的簇,以支持NOMA传输。
- 设计一种深度确定性策略梯度(DDPG)强化学习智能体,实时优化RIS相位移位矩阵。
- 该DDPG智能体通过在长期学习框架中持续探索与利用,学习动态策略,无需预训练模型。
- 该方法将无线环境建模为马尔可夫决策过程,其中智能体根据当前信道状态选择动作(相位移位)。
- 价值函数近似采用随机梯度下降(SGD)方法,以最小化预测值函数与真实值函数之间的均方误差(MSE)。
- 通过迭代更新与基于梯度的优化,该算法确保收敛至损失函数的全局最小值。
实验结果
研究问题
- RQ1RIS辅助的NOMA框架在总数据速率方面是否显著优于传统的RIS辅助OMA?
- RQ2通过强化学习实现的动态RIS相位移位是否优于随机或固定相位移位策略?
- RQ3系统参数中——相位移位粒度或反射单元数量——哪一个对性能提升的影响更大?
- RQ4长期自适应学习模型是否能有效优化用户分组与相位移位,而无需预先训练数据?
- RQ5DRL与RIS及NOMA的集成如何影响系统在动态环境中的可扩展性与适应性?
主要发现
- 所提出的RIS辅助NOMA框架在总数据速率方面优于传统OMA,证明了联合RIS与NOMA优化的有效性。
- 基于DDPG的学习模型成功随时间学习到动态资源分配策略,实现对信道条件变化的持续适应。
- 提高RIS相位移位的粒度可带来可测量的总数据速率增益。
- 降低相位移位粒度和增加反射单元数量,均为提升移动用户总数据速率的有效策略。
- 所提方法收敛至损失函数的全局最小值,确保了长期稳定且最优的策略学习。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。