[Paper Review] Single chip photonic deep neural network with accelerated training
Demonstrates a fully integrated coherent optical DNN on a single chip with in situ training, achieving 92.7% test accuracy on vowel classification and enabling nanosecond inference with ultra-low energy per operation.
As deep neural networks (DNNs) revolutionize machine learning, energy consumption and throughput are emerging as fundamental limitations of CMOS electronics. This has motivated a search for new hardware architectures optimized for artificial intelligence, such as electronic systolic arrays, memristor crossbar arrays, and optical accelerators. Optical systems can perform linear matrix operations at exceptionally high rate and efficiency, motivating recent demonstrations of low latency linear algebra and optical energy consumption below a photon per multiply-accumulate operation. However, demonstrating systems that co-integrate both linear and nonlinear processing units in a single chip remains a central challenge. Here we introduce such a system in a scalable photonic integrated circuit (PIC), enabled by several key advances: (i) high-bandwidth and low-power programmable nonlinear optical function units (NOFUs); (ii) coherent matrix multiplication units (CMXUs); and (iii) in situ training with optical acceleration. We experimentally demonstrate this fully-integrated coherent optical neural network (FICONN) architecture for a 3-layer DNN comprising 12 NOFUs and three CMXUs operating in the telecom C-band. Using in situ training on a vowel classification task, the FICONN achieves 92.7% accuracy on a test set, which is identical to the accuracy obtained on a digital computer with the same number of weights. This work lends experimental evidence to theoretical proposals for in situ training, unlocking orders of magnitude improvements in the throughput of training data. Moreover, the FICONN opens the path to inference at nanosecond latency and femtojoule per operation energy efficiency.
Motivation & Objective
- Motivates energy and throughput limitations of CMOS for deep learning and seeks a scalable photonic solution.
- Proposes a fully integrated photonic circuit with programmable nonlinear optical function units and coherent matrix multiplication units.
- Demonstrates in situ training of a multi-layer photonic DNN on hardware using directional derivatives.
- Shows that optical domain inference can be done without electrical readout between layers and evaluates energy/throughput.
- Provides a path toward real-time learning and ultra-low latency AI hardware on a chip.
Proposed method
- Develops three key components: (i) high-bandwidth programmable nonlinear optical function units (NOFUs); (ii) coherent matrix multiplication units (CMXUs) implemented with a Mach-Zehnder interferometer mesh; (iii) in situ, optically accelerated training computing derivatives on hardware.
- Integrates NOFUs and CMXUs on a single silicon photonic integrated circuit to perform multi-layer DNN operations coherently in the optical domain.
- Uses an in situ training approach where model parameters are updated by measuring directional derivatives along random directions in parameter space, enabling gradient-descent-like optimization without backpropagation.
- Reads out the final optical DNN outputs with an integrated coherent receiver that homodynes the output field with a local oscillator.
- Achieves 92.7% test accuracy on vowel classification, matching the digital model with the same number of weights, using 132 tunable on-chip parameters at 16-bit precision.
Experimental results
Research questions
- RQ1Can a fully integrated coherent optical neural network perform both inference and in situ training on a single chip?
- RQ2What are the achievable accuracies and energy-throughput metrics for a photonic DNN with NOFUs and CMXUs implemented in a commercial silicon photonics process?
- RQ3Does in situ training on hardware using directional derivatives converge to a local minimum for multi-layer photonic networks?
- RQ4How does the on-chip training compare to digital training in terms of final accuracy and training dynamics?
Key findings
- A 3-layer FICONN with 12 NOFUs and three CMXUs operating in telecom C-band demonstrates in situ training and reaches 92.7% test accuracy, identical to a digital model with the same weights.
- The CMXU uses a Mach-Zehnder interferometer mesh to implement a 6×6 unitary with high fidelity (average 0.987 ± 0.007 after error correction).
- The NOFU enables programmable nonlinear activation by detuning a microring resonator via a pn-doped photodiode, achieving ~30 fJ per nonlinear operation and eliminating the need for off-chip amplifiers.
- In situ training computes directional derivatives along random directions in parameter space, updating weights to follow the direction of steepest descent on average, and converges to a local minimum.
- End-to-end on-chip inference incurs an end-to-end loss of 10 dB with per-component insertion losses below 0.1 dB, enabling single-shot inference across all layers without re-amplification.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.