Skip to main content
QUICK REVIEW

[论文解读] Millimeter Wave Communications with an Intelligent Reflector: Performance Optimization and Distributional Reinforcement Learning

Qianqian Zhang, Walid Saad|arXiv (Cornell University)|Feb 24, 2020
Advanced Wireless Communication Technologies参考文献 34被引用 20
一句话总结

本文提出了一种基于可重构智能反射面(IR)的毫米波通信联合优化框架,针对完美信道状态信息(CSI)采用迭代预编码与反射波束成形,针对非完美CSI则采用分位数回归分布强化学习(QR-DRL)方法。QR-DRL方法实现了稳定的收敛性,并在有限CSI条件下相比Q-learning将下行链路速率提升了10%,与固定IR方案相比总速率增益超过30%。

ABSTRACT

In this paper, a novel framework is proposed to optimize the downlink multi-user communication of a millimeter wave base station, which is assisted by a reconfigurable intelligent reflector (IR). In particular, a channel estimation approach is developed to measure the channel state information (CSI) in real-time. First, for a perfect CSI scenario, the precoding transmission of the BS and the reflection coefficient of the IR are jointly optimized, via an iterative approach, so as to maximize the sum of downlink rates towards multiple users. Next, in the imperfect CSI scenario, a distributional reinforcement learning (DRL) approach is proposed to learn the optimal IR reflection and maximize the expectation of downlink capacity. In order to model the transmission rate's probability distribution, a learning algorithm, based on quantile regression (QR), is developed, and the proposed QR-DRL method is proved to converge to a stable distribution of downlink transmission rate. Simulation results show that, in the error-free CSI scenario, the proposed approach yields over 30% and 2-fold increase in the downlink sum-rate, compared with a fixed IR reflection scheme and direct transmission scheme, respectively. Simulation results also show that by deploying more IR elements, the downlink sum-rate can be significantly improved. However, as the number of IR components increases, more time is required for channel estimation, and the slope of increase in the IR-aided transmission rate will become smaller. Furthermore, under limited knowledge of CSI, simulation results show that the proposed QR-DRL method, which learns a full distribution of the downlink rate, yields a better prediction accuracy and improves the downlink rate by 10% for online deployments, compared with a Q-learning baseline.

研究动机与目标

  • 为解决在阻塞和非完美信道状态信息(CSI)条件下实现可靠毫米波通信的挑战。
  • 通过联合波束成形与反射控制优化多用户毫米波系统中智能反射面(IR)的下行链路总速率。
  • 在CSI不完美时开发一种鲁棒的基于学习的IR反射优化方法,确保在不确定性下的稳定性能。
  • 利用分布强化学习对下行链路传输速率的完整分布进行建模,以提高预测精度。

提出的方法

  • 提出一种迭代算法,在完美CSI条件下联合优化基站(BS)预编码与IR反射系数,以最大化总速率。
  • 引入一种基于分位数回归(QR)的分布强化学习(DRL)框架,以建模下行链路传输速率的完整分布。
  • 开发一种QR-DRL算法,通过最小化分位数回归损失来学习最优IR反射策略,确保收敛至稳定分布。
  • 通过1-Wasserstein距离与压缩映射理论,证明QR-DRL方法的理论收敛性,其误差界取决于分位数分辨率。
  • 采用基于1-Wasserstein距离的度量方法,衡量价值函数之间的分布距离,从而实现鲁棒的策略学习。
  • 使用投影算子 Π_{W₁} 近似真实传输速率分布,确保算法的稳定性和收敛性。

实验结果

研究问题

  • RQ1在完美CSI条件下,如何在IR辅助的毫米波系统中联合优化预编码与反射波束成形,以最大化下行链路总速率?
  • RQ2与直接传输和固定IR反射方案相比,所提出的优化框架在性能上有哪些提升?
  • RQ3如何利用分布强化学习在非完美CSI条件下学习最优IR反射策略?
  • RQ4在有限CSI场景下,QR-DRL是否能实现比传统Q-learning更高的预测精度和速率性能?
  • RQ5IR单元数量对系统性能及信道估计开销有何影响?

主要发现

  • 在完美CSI场景下,所提方法相比固定IR反射方案,下行链路总速率提升超过30%。
  • 与无IR辅助的直接传输相比,所提方法使总速率提升两倍。
  • 随着IR单元数量的增加,总速率持续提升,但提升速率因信道估计开销增加而逐渐减缓。
  • 在非完美CSI条件下,QR-DRL方法在在线部署中相比Q-learning基线将下行链路速率提升了10%。
  • 理论分析证明,QR-DRL算法收敛至稳定的传输速率分布,且收敛误差随分位数数量增加而减小。
  • 在压缩映射条件下,QR-DRL策略的收敛性得到保证,且误差界随分位数分辨率提高而缩小。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。