Skip to main content
QUICK REVIEW

[论文解读] Noise-Robust End-to-End Quantum Control using Deep Autoregressive Policy Networks

Jiahao Yao, Paul Köttering|arXiv (Cornell University)|Dec 12, 2020
Quantum Computing Algorithms and Architecture参考文献 104被引用 9
一句话总结

本文提出RL-QAOA,一种抗噪声的端到端量子控制框架,通过深度自回归策略网络联合优化连续门持续时间和离散酉矩阵序列。通过结合近端策略优化与混合连续-离散动作空间,该方法在经典噪声和量子噪声下对非可积量子伊辛模型的基态制备任务中表现出优越性能,优于先前的PG-QAOA和CD-QAOA方法,适用于多种噪声模型。

ABSTRACT

Variational quantum eigensolvers have recently received increased attention, as they enable the use of quantum computing devices to find solutions to complex problems, such as the ground energy and ground state of strongly-correlated quantum many-body systems. In many applications, it is the optimization of both continuous and discrete parameters that poses a formidable challenge. Using reinforcement learning (RL), we present a hybrid policy gradient algorithm capable of simultaneously optimizing continuous and discrete degrees of freedom in an uncertainty-resilient way. The hybrid policy is modeled by a deep autoregressive neural network to capture causality. We employ the algorithm to prepare the ground state of the nonintegrable quantum Ising model in a unitary process, parametrized by a generalized quantum approximate optimization ansatz: the RL agent solves the discrete combinatorial problem of constructing the optimal sequences of unitaries out of a predefined set and, at the same time, it optimizes the continuous durations for which these unitaries are applied. We demonstrate the noise-robust features of the agent by considering three sources of uncertainty: classical and quantum measurement noise, and errors in the control unitary durations. Our work exhibits the beneficial synergy between reinforcement learning and quantum control.

研究动机与目标

  • 开发一种统一的强化学习框架,同时优化量子控制协议中的连续参数与离散参数。
  • 增强对多种噪声源的鲁棒性——包括经典测量噪声、量子测量噪声以及门持续时间误差——这些噪声在NISQ设备中普遍存在。
  • 通过在单一可微策略中实现酉矩阵序列顺序与门持续时间的端到端优化,扩展现有基于QAOA的方法。
  • 在现实噪声条件下,证明该方法在制备非可积量子伊辛链基态方面的有效性。

提出的方法

  • 该方法采用深度自回归神经网络建模一种混合策略,以捕捉序列量子控制协议中动作之间的因果依赖关系。
  • 使用针对联合连续与离散动作空间改进的近端策略优化(PPO)算法进行策略训练。
  • 连续动作通过紧支撑的Beta分布参数化,以确保门持续时间的有界性。
  • 离散动作对应于从预定义集合中选择酉矩阵,其序列顺序通过端到端方式优化。
  • 该算法应用于广义QAOA量子电路,支持可变的酉操作序列与优化的门持续时间。
  • 在三种不确定性模型下评估抗噪声性能:经典测量噪声、量子测量噪声以及门持续时间的随机误差。

实验结果

研究问题

  • RQ1单一强化学习智能体能否有效优化量子控制中的连续门持续时间与离散酉矩阵序列?
  • RQ2所提出的混合策略网络在NISQ设备中常见的现实噪声源下表现如何?
  • RQ3策略的自回归结构是否能实现更优的因果建模并提升量子控制中的收敛性能?
  • RQ4在噪声环境下,RL-QAOA与PG-QAOA和CD-QAOA相比,在鲁棒性与性能方面表现如何?
  • RQ5该方法是否能在无需架构重构的情况下泛化至可变长度的控制序列?

主要发现

  • RL-QAOA在高度非绝热区域成功制备了非可积量子伊辛模型的基态,在所有测试的噪声模型下均优于PG-QAOA和CD-QAOA。
  • 该方法展现出对噪声的无偏鲁棒性,无论不确定性来源是经典测量噪声、量子测量噪声还是门持续时间误差,均能保持高性能。
  • 深度自回归策略网络有效建模了时间因果性,实现了对序列量子控制协议的稳定高效优化。
  • 基于混合PPO的训练框架在同时存在连续与离散动作空间时实现了稳定学习,克服了先前基于梯度的优化器在噪声环境下的局限性。
  • 即使总协议时长不固定,该算法依然有效,因为引入了'停止'动作,支持可变长度序列。
  • 数值模拟结果证实,在所有考虑的噪声条件下,RL-QAOA的基态保真度均高于基线方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。