[论文解读] Real-Time FJ/MAC PDE Solvers via Tensorized, Back-Propagation-Free Optical PINN Training
该论文提出了一种首个基于张量化光子神经网络(TONN)与相位域调谐的片上、无反向传播(BP-free)光学训练框架,用于物理信息神经网络(PINNs),实现了高维偏微分方程(PDE)的实时、超低能耗求解(1.36 J,1.15 s)——在20维汉密尔顿-雅可比-贝尔曼(HJB)PDE上实现,MZI数量减少1,170倍,光子每MAC能量效率达fJ级别。
Solving partial differential equations (PDEs) numerically often requires huge computing time, energy cost, and hardware resources in practical applications. This has limited their applications in many scenarios (e.g., autonomous systems, supersonic flows) that have a limited energy budget and require near real-time response. Leveraging optical computing, this paper develops an on-chip training framework for physics-informed neural networks (PINNs), aiming to solve high-dimensional PDEs with fJ/MAC photonic power consumption and ultra-low latency. Despite the ultra-high speed of optical neural networks, training a PINN on an optical chip is hard due to (1) the large size of photonic devices, and (2) the lack of scalable optical memory devices to store the intermediate results of back-propagation (BP). To enable realistic optical PINN training, this paper presents a scalable method to avoid the BP process. We also employ a tensor-compressed approach to improve the convergence and scalability of our optical PINN training. This training framework is designed with tensorized optical neural networks (TONN) for scalable inference acceleration and MZI phase-domain tuning for extit{in-situ} optimization. Our simulation results of a 20-dim HJB PDE show that our photonic accelerator can reduce the number of MZIs by a factor of $1.17 imes 10^3$, with only $1.36$ J and $1.15$ s to solve this equation. This is the first real-size optical PINN training framework that can be applied to solve high-dimensional PDEs.
研究动机与目标
- 为自主系统和医学成像等能效受限应用场景提供实时、低能耗的PDE求解能力。
- 克服光学神经网络中因光子器件尺寸过大及缺乏可扩展光存储而造成的反向传播硬件不兼容问题。
- 开发一种可扩展、鲁棒且能效高效的无BP梯度估计与张量压缩结合的光学PINN训练框架。
- 在集成光子平台上实现大规模PINNs的片上训练可行性验证。
提出的方法
- 提出一种无反向传播的训练方法,仅通过额外推理估计梯度与导数,避免误差反馈及硬件不友好的BP过程。
- 采用张量列车(TT)分解压缩权重矩阵,减少MZI数量,提升收敛性与可扩展性。
- 采用基于MZI的相位域调谐的张量化光子神经网络(TONN),实现原位优化与可扩展推理加速。
- 通过直接在制造后的光子器件上进行硬件感知调优,提升对硬件非理想性的鲁棒性。
- 采用多周期推理架构(TONN-2中为64周期),降低面积与插入损耗,实现高效的光子计算。
- 在紧凑面积内集成光子元件,包括混合硅激光器、微环调制器、MZI光栅阵列及光电二极管,实现片上集成。
实验结果
研究问题
- RQ1能否设计一种无反向传播的光学训练框架,以实现在光子芯片上大规模PINNs的可扩展、高效训练?
- RQ2张量压缩在光学PINN训练中如何减少MZI数量并改善收敛性?
- RQ3在存在硬件非理想性的情况下,片上基于推理的梯度估计在性能上是否显著优于软件仿真或片外训练?
- RQ4在TONN架构中采用多周期推理进行PDE求解时,能量与延迟之间存在何种权衡?
- RQ5所提出的框架能否在20D HJB等高维PDE上实现实时、fJ/MAC级别的能效?
主要发现
- 所提出的光学PINN训练框架在20D HJB PDE上相比传统ONN将MZI数量减少了1.17×10³倍。
- 系统在求解20D HJB PDE时总能耗仅为1.36 J,总延迟为1.15 s,实现真正意义上的实时性能。
- TONN-2的光子芯片面积为26 mm²,显著小于TONN-1的648 mm²,尽管因64周期推理导致计算延迟更高。
- 采用额外推理进行梯度估计的无BP方法在硬件非理想性下表现出优于软件仿真或片外训练的鲁棒性。
- 该框架支持最大1024×1024的全连接网络,证明了其在大规模PINNs中的可扩展性。
- TONN-2中单次推理能耗降低至5.05×10⁻⁹ J,实现fJ/MAC级别的光子能效。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。