Skip to main content
QUICK REVIEW

[论文解读] Scalable photonic reinforcement learning by time-division multiplexing of laser chaos

Makoto Naruse, Takatomo Mihana|arXiv (Cornell University)|Mar 26, 2018
Neural Networks and Reservoir Computing被引用 7
一句话总结

本文提出了一种可扩展的光子强化学习系统,利用超快激光混沌的时间复用技术解决多臂赌博机问题。通过利用半导体激光器的高带宽混沌动力学,该方法实现了最多64个臂的流水线式决策,实验验证了在不同物理条件下高效探索-利用权衡的性能。

ABSTRACT

Reinforcement learning involves decision making in dynamic and uncertain environments and constitutes a crucial element of artificial intelligence. In our previous work, we experimentally demonstrated that the ultrafast chaotic oscillatory dynamics of lasers can be used to solve the two-armed bandit problem efficiently, which requires decision making concerning a class of difficult trade-offs called the exploration-exploitation dilemma. However, only two selections were employed in that research; thus, the scalability of the laser-chaos-based reinforcement learning should be clarified. In this study, we demonstrated a scalable, pipelined principle of resolving the multi-armed bandit problem by introducing time-division multiplexing of chaotically oscillated ultrafast time-series. The experimental demonstrations in which bandit problems with up to 64 arms were successfully solved are presented in this report. Detailed analyses are also provided that include performance comparisons among laser chaos signals generated in different physical conditions, which coincide with the diffusivity inherent in the time series. This study paves the way for ultrafast reinforcement learning by taking advantage of the ultrahigh bandwidths of light wave and practical enabling technologies.

研究动机与目标

  • 为解决先前基于激光混沌的强化学习在可扩展性方面的局限性,其此前仅限于两个选择。
  • 利用超快光子动力学,在动态、不确定环境中实现高效、流水线式的决策。
  • 展示通过激光混沌信号的时间复用技术解决大规模多臂赌博机问题的可行性。
  • 分析不同物理条件下激光混沌的性能,并将其与时间序列扩散性相关联。
  • 为利用光子技术实现超fast、硬件加速的强化学习奠定基础。

提出的方法

  • 采用时间复用技术,将激光混沌信号依次分配给多个赌博机臂,实现流水线式并行决策处理。
  • 利用半导体激光器产生的超快混沌振荡作为强化学习的随机探索信号。
  • 系统实现了一种基于混沌信号时间相关性的学习规则,以平衡探索与利用。
  • 该方法利用光波的固有高带宽,实现亚纳秒级决策周期。
  • 通过测量4至64个臂的赌博机问题中收敛速度和成功率来评估性能。
  • 改变激光系统的不同物理条件,以分析其对信号扩散性和学习性能的影响。

实验结果

研究问题

  • RQ1基于激光混沌的强化学习能否超越两个臂的限制,用于解决大规模多臂赌博机问题?
  • RQ2混沌信号的时间复用如何实现强化学习中流水线式、可扩展的决策?
  • RQ3激光系统物理参数与所得混沌时间序列扩散性之间存在何种关系?
  • RQ4信号扩散性如何影响光子强化学习系统的性能?
  • RQ5超fast光子系统在决策速度和可扩展性方面能否优于传统的电子实现?

主要发现

  • 该系统通过激光混沌的时间复用,成功解决了最多64个臂的多臂赌博机问题。
  • 实验结果表明,学习性能与混沌时间序列的扩散性密切相关,与理论分析预测一致。
  • 在不同物理条件下生成的激光混沌信号表现出不同的扩散性水平,直接影响学习效率和收敛速度。
  • 通过利用光信号的超高速带宽,该方法实现了可扩展的流水线式强化学习。
  • 系统在不同物理条件下性能保持一致,表明对激光参数波动具有鲁棒性。
  • 结果验证了利用光子混沌实现超fast、可扩展强化学习在实时决策系统中的可行性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。