[Paper Review] Self-Correcting Quantum Many-Body Control using Reinforcement Learning with Tensor Networks
This paper introduces a reinforcement learning framework using matrix product states (MPS) to enable efficient, scalable control of quantum many-body systems. By representing both the quantum state and the RL agent's policy via tensor networks (QMPS), the method achieves linear scaling with system size, enabling control of larger 1D systems than previously possible with pure neural networks, while demonstrating robustness to noise and self-correction in dynamic environments.
Quantum many-body control is a central milestone en route to harnessing quantum technologies. However, the exponential growth of the Hilbert space dimension with the number of qubits makes it challenging to classically simulate quantum many-body systems and consequently, to devise reliable and robust optimal control protocols. Here, we present a novel framework for efficiently controlling quantum many-body systems based on reinforcement learning (RL). We tackle the quantum control problem by leveraging matrix product states (i) for representing the many-body state and, (ii) as part of the trainable machine learning architecture for our RL agent. The framework is applied to prepare ground states of the quantum Ising chain, including states in the critical region. It allows us to control systems far larger than neural-network-only architectures permit, while retaining the advantages of deep learning algorithms, such as generalizability and trainable robustness to noise. In particular, we demonstrate that RL agents are capable of finding universal controls, of learning how to optimally steer previously unseen many-body states, and of adapting control protocols on-the-fly when the quantum dynamics is subject to stochastic perturbations. Furthermore, we map the QMPS framework to a hybrid quantum-classical algorithm that can be performed on noisy intermediate-scale quantum devices and test it under the presence of experimentally relevant sources of noise.
Motivation & Objective
- To address the exponential scaling challenge in simulating and controlling quantum many-body systems due to Hilbert space growth.
- To develop a reinforcement learning agent capable of controlling large 1D quantum systems beyond the reach of exact diagonalization.
- To integrate matrix product states (MPS) into the RL architecture to enable efficient, scalable policy learning.
- To demonstrate robustness of the control protocol to stochastic noise and coherent gate errors.
- To map the framework to a hybrid quantum-classical algorithm suitable for NISQ devices.
Proposed method
- The quantum state is represented using matrix product states (MPS), enabling efficient simulation of 1D systems with area-law entanglement.
- The RL agent's policy is implemented as a hybrid MPS-neural network architecture (QMPS), where the MPS structure encodes the state-action value function.
- The framework uses policy-gradient reinforcement learning with Q-values computed via a differentiable MPS representation, allowing backpropagation through the tensor network.
- The method enables training on systems with up to 128 qubits, far exceeding typical deep RL limits for quantum control.
- A hybrid quantum-classical implementation is designed, where quantum circuits sample fidelity and gradients, while classical networks update the policy.
- The framework is tested under realistic noise models, including depolarizing, amplitude/phase damping, and coherent gate errors.
Experimental results
Research questions
- RQ1Can a reinforcement learning agent trained on matrix product states control larger quantum many-body systems than traditional deep learning approaches?
- RQ2Can the QMPS architecture maintain high fidelity and generalization when preparing ground states across different phases of the Ising model?
- RQ3How does the RL agent adapt to and self-correct under stochastic perturbations or noise in the quantum dynamics?
- RQ4To what extent can the QMPS framework be implemented on noisy intermediate-scale quantum (NISQ) devices with realistic noise models?
- RQ5Does the use of tensor networks in the RL policy architecture improve robustness to noise compared to standard neural networks?
Key findings
- The QMPS framework successfully prepares ground states of the 1D transverse and mixed-field Ising models, including in the critical region, with fidelity exceeding 0.99 for systems up to 128 qubits.
- The method achieves near-unity success rates (close to 100%) with as few as 500 measurement shots per fidelity evaluation, demonstrating robustness to sampling noise.
- The agent maintains high performance under amplitude and phase damping noise, achieving success rates near unity for error rates below 10⁻³.
- The QMPS agent self-corrects protocols under coherent gate errors, maintaining success rates for standard deviation σ < 0.5 in rotation angle noise.
- The truncated QMPS with bond dimension χ=2 performs significantly worse than χ=4, confirming the necessity of sufficient entanglement capacity in the tensor network.
- The hybrid quantum-classical implementation enables feasible training on NISQ devices, with performance converging within 10⁴ shots under ideal conditions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.