Skip to main content
QUICK REVIEW

[论文解读] Laziness, Barren Plateau, and Noise in Machine Learning

Junyu Liu, Zexi Lin|arXiv (Cornell University)|Jun 19, 2022
Quantum Computing Algorithms and Architecture参考文献 57被引用 12
一句话总结

本文将'迟滞'(laziness)定义为变分量子与经典神经网络中一种独立的现象,其特征是由于过参数化导致参数更新被抑制——这与损失函数变化可忽略的' barren plateaus'( barren plateaus)现象相区别。研究发现,在过参数化条件下,量子线路对噪声具有鲁棒性,且能避免陷入 barren plateaus,理论与数值证据均表明在过参数化及最优学习率下具备噪声鲁棒性。

ABSTRACT

We define \emph{laziness} to describe a large suppression of variational parameter updates for neural networks, classical or quantum. In the quantum case, the suppression is exponential in the number of qubits for randomized variational quantum circuits. We discuss the difference between laziness and \emph{barren plateau} in quantum machine learning created by quantum physicists in \cite{mcclean2018barren} for the flatness of the loss function landscape during gradient descent. We address a novel theoretical understanding of those two phenomena in light of the theory of neural tangent kernels. For noiseless quantum circuits, without the measurement noise, the loss function landscape is complicated in the overparametrized regime with a large number of trainable variational angles. Instead, around a random starting point in optimization, there are large numbers of local minima that are good enough and could minimize the mean square loss function, where we still have quantum laziness, but we do not have barren plateaus. However, the complicated landscape is not visible within a limited number of iterations, and low precision in quantum control and quantum sensing. Moreover, we look at the effect of noises during optimization by assuming intuitive noise models, and show that variational quantum algorithms are noise-resilient in the overparametrization regime. Our work precisely reformulates the quantum barren plateau statement towards a precision statement and justifies the statement in certain noise models, injects new hope toward near-term variational quantum algorithms, and provides theoretical connections toward classical machine learning. Our paper provides conceptual perspectives about quantum barren plateaus, together with discussions about the gradient descent dynamics in \cite{together}.

研究动机与目标

  • 阐明变分量子与经典神经网络中'迟滞'(参数更新被抑制)与' barren plateaus'(损失函数变化被抑制)之间的区别。
  • 基于神经正切核(NTK)理论,建立理论框架,解释为何过参数化量子线路尽管损失景观复杂,仍能避免陷入 barren plateaus。
  • 研究在过参数化条件下,变分量子算法对噪声的鲁棒性,特别是针对现实噪声模型的响应。
  • 建立量子 barren plateaus 与优化精度极限之间的精确联系,将问题重新表述为噪声鲁棒性声明。
  • 通过随机化硬件高效ansatz的数值模拟验证理论预测,结果显示残差误差与噪声标度的理论结果与数值结果高度一致。

提出的方法

  • 将'迟滞'定义为由于高希尔伯特空间维度(量子)或网络宽度(经典)导致的参数更新抑制(δθμ),并将其与 barren plateau(δℒ ≪ 1)相区分。
  • 应用神经正切核(NTK)理论分析过参数化量子与经典网络中的梯度动力学,表明在过参数化下 dQNTK 实现集中化。
  • 推导出在噪声作用下参数演化过程的随机微分方程(SDE)模型,其中残差误差 ε(t) 满足 ε(t) ≈ (1−ηK)^t ε(0) + 噪声项。
  • 通过将噪声引起的误差与无噪声残差误差相等,提出噪声鲁棒性条件,推导出噪声主导残差误差的临界时间 T_noise。
  • 利用条件 η ≈ O(1/K) 最小化最终残差误差,表明其与 dQNTK 集中化条件一致,前提是 ε(0) < O(L√N)。
  • 在含 4 量子比特、L=64 的随机化硬件高效ansatz上进行数值模拟,将平均残差误差与标准差与理论预测进行对比。

实验结果

研究问题

  • RQ1在变分量子线路中,'迟滞'与 barren plateaus 现象有何不同?
  • RQ2为何过参数化经典与量子神经网络尽管参数更新被抑制,仍能保持实用性?
  • RQ3在存在噪声的情况下,变分量子算法在何种条件下可避免陷入 barren plateaus?
  • RQ4噪声如何影响过参数化变分量子线路的收敛性与残差误差?
  • RQ5能否将量子 barren plateau 问题重新解释为精度极限?在何种噪声模型下该解释成立?

主要发现

  • 迟滞——即参数更新被抑制——在过参数化经典与量子网络中均存在,但并不必然导致 barren plateaus。
  • 在过参数化量子线路中,损失景观复杂且存在大量局部极小值,但系统因噪声鲁棒性与 dQNTK 集中化而避免陷入 barren plateaus。
  • 数值模拟显示,残差误差随噪声标准差(σθ)与学习率(η)变化的理论预测与实验结果高度一致。
  • 最优学习率被确定为 η ≈ O(1/K),该值可最小化最终残差误差,并与 dQNTK 集中化条件保持一致。
  • 临界时间 T_noise(噪声主导残差误差)已通过解析推导并经数值验证,表明噪声影响仅在大量迭代后才变得显著。
  • 最终残差误差 σ_ε(∞) 的标度关系为 √(2/π) ⋅ σ_θ / √(2η − η²K),在数值实验中与理论预测在 90% 置信区间内一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。