[Paper Review] Training neural networks with end-to-end optical backpropagation
This paper presents the first complete end-to-end optical backpropagation framework for training neural networks using saturable absorbers in a rubidium vapor cell. By leveraging a pump-probe configuration where forward signals act as pumps and backward signals as probes, the method enables all-optical gradient computation through nonlinear activation layers, achieving superior performance over in silico training with minimal hardware overhead.
Optics is an exciting route for the next generation of computing hardware for machine learning, promising several orders of magnitude enhancement in both computational speed and energy efficiency. However, to reach the full capacity of an optical neural network it is necessary that the computing not only for the inference, but also for the training be implemented optically. The primary algorithm for training a neural network is backpropagation, in which the calculation is performed in the order opposite to the information flow for inference. While straightforward in a digital computer, optical implementation of backpropagation has so far remained elusive, particularly because of the conflicting requirements for the optical element that implements the nonlinear activation function. In this work, we address this challenge for the first time with a surprisingly simple and generic scheme. Saturable absorbers are employed for the role of the activation units, and the required properties are achieved through a pump-probe process, in which the forward propagating signal acts as the pump and backward as the probe. Our approach is adaptable to various analog platforms, materials, and network structures, and it demonstrates the possibility of constructing neural networks entirely reliant on analog optical processes for both training and inference tasks.
Motivation & Objective
- To overcome the longstanding challenge of implementing backpropagation entirely in optics, particularly through nonlinear activation functions.
- To eliminate the need for digital computation or opto-electronic conversions during training in analog optical neural networks.
- To enable real-time, in-situ training of optical neural networks using only optical components and processes.
- To demonstrate that optical backpropagation can outperform conventional in silico training due to reduced 'reality gap' effects.
- To develop a generic, adaptable framework applicable to various optical platforms, materials, and network architectures.
Proposed method
- Utilizes a pump-probe process in a rubidium vapor cell, where the forward signal (pump) saturates atomic transitions and the backward signal (probe) measures the derivative of the activation function.
- Employs saturable absorbers as nonlinear activation units, with the probe beam's transmission modulated by the pump intensity to emulate the derivative of the activation function.
- Applies optical matrix-vector multiplication (MVM) using spatial light modulators (SLMs) and cylindrical lenses for 'fan-in' and 'fan-out' operations to implement linear layers.
- Performs backward pass via reverse beam propagation through the same optical setup, with DMDs and SLMs implementing transposed weight matrices and error backpropagation.
- Uses three measurements per training iteration (pump only, probe only, both) to subtract background offsets from fluorescence and unabsorbed probe, ensuring accurate gradient estimation.
- Employs coherent detection with high-speed cameras to measure activation and error vectors, which are used to compute weight gradients and update SLMs in real time.

Experimental results
Research questions
- RQ1Can backpropagation through nonlinear activation functions be implemented entirely in optics without digital intervention?
- RQ2Can a single optical system perform both inference and training using reverse signal propagation?
- RQ3Does all-optical backpropagation outperform traditional in silico training in terms of accuracy and robustness to hardware imperfections?
- RQ4Can the pump-probe configuration in a saturable absorber accurately emulate the derivative of a nonlinear activation function during backward pass?
- RQ5Is the proposed method generalizable across different optical platforms, materials, and neural network architectures?
Key findings
- The optical backpropagation framework successfully computes gradients through nonlinear activation layers using saturable absorption in a rubidium vapor cell.
- The method achieves accurate gradient estimation by subtracting background signals through three-measurement calibration, compensating for fluorescence and residual probe transmission.
- The network trained via all-optical backpropagation outperforms networks trained using conventional in silico methods, demonstrating reduced error from the 'reality gap'.
- The system operates with minimal hardware overhead, using standard components like SLMs, DMDs, and a single vapor cell for both forward and backward passes.
- The approach is generic and adaptable to various optical platforms, materials, and network structures, including those based on interferometric or other linear layer implementations.
- The experimental results confirm that end-to-end optical training is feasible and effective, marking a critical step toward fully analog optical machine learning.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.